Compare commits

...
21 Commits
Author SHA1 Message Date
YoursFunny 9e873131d4 chore: bump version to 1.8.0 2026-09-20 17:41:20 +08:00
YoursFunny 5b77d14497 fix(commands): make /debug preview the caption a link would actually send
`/debug` reported `Fetched::caption`, the site's built-in caption, so a
chat's `/set_format` override never showed up in the preview — the command
looked like a no-op, and the `/set_format` success reply now tells users to
preview with `/debug`, which made that advice wrong.

`preview_caption` mirrors the send paths instead: the per-site format
override (`caption_from_fields`, empty → built-in) plus the long-post
quoting. Reported from a live run against the test bot.
2026-09-20 17:30:55 +08:00
YoursFunny d4c36feb9a fix(commands): make /set_format and /clear_cache actually parse
teloxide's `split` parser takes exactly one space-separated token per
field, so `/set_format <site> <format>` — a two-token command — never
parsed: `Command::parse` failed, `message_handler` fell through to the URL
flow, and the user got silence. `/clear_cache` without its optional link
failed the same way ("too few arguments"), so clearing everything was
unreachable. Both now use the crate's `parse_arg_remainder` (whole
remainder, trimmed), which is what their executors were already written
against (`split_once(char::is_whitespace)`).

Found by driving the real binary against a scripted fake Bot API: the
`/set_format` replies never appeared while `/set_template` and the other
single-token commands did. `every_documented_invocation_parses` now pins
every documented form, which is what should have caught it.

The placeholder validation and `-` reset added earlier only work now that
the command reaches its executor at all.
2026-09-20 16:42:40 +08:00
YoursFunny 5d0acdac01 feat(ux): onboard users, expose the chat's settings, name failed posts
`/start` was "Hello!" and `/help` was the bare command list teloxide can
render — no argument syntax, no caption placeholders, no mention that
links only work in private chats. Both now carry that guidance, and the
bot's profile description / short description are set at startup so a
shared link says what the bot does.

`/settings` reports what this chat is configured to do (forward channel,
edit-before-forward, per-site formats, saved templates) to anyone in the
chat — `/bot_dict` is a raw admin-only dump. Templates can be removed
(`/remove_template`, listing the live names on a typo) and the prompt's
keyboard folds 3 per row with a cap: Telegram rejects a keyboard over 100
buttons outright, which would silently drop the whole prompt.

Inline results hand URLs to Telegram, which fetches them without any
site headers — pixiv's pximg.net answers 403 to that, so those items are
skipped instead of shipped broken. `needs_media_headers` answers that
question from the same per-site rule the downloader uses.

Dead-letter and retry notices name the failing post and the cause
(`failure_text`), since "Task failed after retries: task failed after 2
retries" said neither which link it was nor what happened.
2026-09-20 15:57:45 +08:00
YoursFunny d6707133cc feat(ux): answer every link, name fetch failures, keep the chat action alive
Four ways a user could get silence are closed: a registered-but-disabled
site (pixiv without a token) now answers instead of being dropped as an
unsupported link, `/test` on such a link replies instead of doing nothing,
a supported link posted in a group gets a one-line hint (channels stay
silent), and fetch failures name their cause — gone / withheld / source
risk control / site disabled / source down — instead of one generic
sentence. `FetchError::Disabled` carries the "matched but switched off"
answer, which `find_site` used to fold into `Ok(None)`.

A withheld tweet no longer degrades to "no media": without
`TWITTER_AUTH_TOKEN` it stays `Sensitive` so the reply says the media is
age-restricted, and a failed authenticated fallback propagates its own
class instead of masquerading as an empty post (`empty_fetched` is gone).

Long jobs stop looking stalled: `run_with_chat_action` re-sends the chat
action every 4s while the pipeline is pending and the hint switches from
typing to send-photo/video once the media kinds are known. Media groups
go from 9 to Telegram's 10.

`/set_format` rejects unknown `{…}` placeholders (a typo used to be
published verbatim in every caption) and resets with `-`. The
edit-before-forward prompt states its TTL and that Confirm is required,
gains a Skip button, and is rewritten in place to "expired" by the sweep
— an edit, never a new message, so a background timer cannot wake a chat.
2026-09-20 15:48:46 +08:00
YoursFunny fb601f4d5d chore: bump version to 1.7.0 2026-09-18 01:17:45 +08:00
YoursFunny d8dd4fa91e test(cache): pin that pre-split entries still parse
A payload written before the title/content split has no `content` field;
`#[serde(default)]` is what keeps it readable, and the cache deletes any
payload it cannot parse — so dropping that default would silently evict
entries rather than degrade them. The test inserts the literal pre-split
JSON and asserts it comes back with its text left in `title` (no
migration: the entry lives one TTL and moving the text would only
reshuffle `/set_format` placeholders until it expires) and its stored
caption untouched.
2026-09-18 01:14:09 +08:00
YoursFunny af96caff40 feat(send): quote a long post's text in an expandable blockquote
A post whose text (the split `title` plus `content`, joined by
`site::compose_text`) reaches `CAPTION_QUOTE_TEXT_CHARS` — default 200,
`0` disables — now has that text wrapped in `<blockquote expandable>`
inside its caption, leaving the URL and author line outside the quote.

Applied at the send boundary (`send_media_sequence`, `send_animation` and
the inline answers), where the caption is already truncated and the same
cache snapshot supplies the text, so a fresh send, a link-cache resend
and a queued retry all decide identically. The text is located as what
follows the author link, with the visible prefix accepted as a match
because `truncate_caption` may cut inside it — that keeps the longest
posts, the ones that most need folding, quoted. Captions whose layout
moves the text elsewhere (pixiv's title-inside-a-link, a `/set_format`
that puts `{title}`/`{content}` first) stay unquoted rather than risking
a blockquote nested in a tag, and a caption that already carries one is
never wrapped again.

Telegram measures a caption *after entities parsing*, so the tags cost no
length and the 1024-character limit cannot be breached; retries replay
the unwrapped caption, so a threshold change takes effect immediately.
The edit-before-forward rewrite stays unquoted by design.

Verified against Telegram: a media-group caption built this way comes
back with `caption_entities` `url` @0, `text_link` @50,
`expandable_blockquote` @56 — the quote starts after the author line.
2026-09-18 00:50:30 +08:00
YoursFunny 52184ba6fb refactor(x-media): split the post title from its content
`Fetched.title` carried whatever text the platform had — a tweet's body,
a bilibili dynamic's body, a pixiv artwork's title — which was enough
while x/twitter (no title at all) set the shape. The platforms actually
disagree: pixiv has a title *and* a description, bilibili has an opus
headline *and* a body. Posts now carry both:

- `title`: the platform's title (a pixiv artwork title, a bilibili opus
  headline or video card title), empty on text-only platforms;
- `content`: the body (tweet / bsky / misskey text, bilibili dynamic
  body, and pixiv's description — fetched for the first time here and
  flattened from the app API's HTML to plain text).

`{content}` joins the caption-format placeholders, so a custom
`/set_format` can include a pixiv description. The built-in captions keep
producing byte-identical output: `compose_text` joins the two fields the
same way the single field already was, and bilibili's forward marker
(`//@author:`) now lands in `content` behind the head line's `title`.

`CachedPost.content` is `#[serde(default)]`, so link-cache entries and
queued task payloads written before the split still parse, their text
living in `title`.
2026-09-18 00:44:49 +08:00
YoursFunny 0eb4e5c78d fix(sites): request the opus serialization so bilibili posts keep their text
An image/text post fetched without `features=itemOpusStyle` comes back in
bilibili's legacy shape, where the post's body and headline are gone
completely — `desc: null`, no `major.opus` — so `title` (and `{title}`)
stayed empty for exactly the posts that do have content
(`opus/1248857553488576532`: legacy `desc: null`, flagged
`major.opus.summary.text = "[doge_金箍]黑白搭配"`). The same flag also
moves the pictures to `major.opus.pics` (key `url`, not `src`).

Text now falls back opus (headline + body) → `desc.text` → archive card
title; media falls back `opus.pics` → `draw.items` → archive cover, so
the legacy shapes keep working if the flag is ever retired.

Verified live: the reported link now yields title "[doge_金箍]黑白搭配"
with its picture; AV dynamics keep their card title; forwards keep
`desc.text` and gain `//@` composition unchanged.
2026-09-17 22:34:11 +08:00
YoursFunny 5c51de217a fix(sites): use the archive title when a bilibili dynamic has no body
A 视频投稿动态 (`DYNAMIC_TYPE_AV`) carries no body at all: the API
answers `desc: null` and the content is the archive card, so `title`
(and `{title}` in caption formats) stayed empty for the most common
dynamic type. Audited 24 live dynamics: every dynamic that *has* text
(a 图文 post, a forward, a text post) keeps it in
`module_dynamic.desc.text` — only the AV card has none, so the video
title now stands in, mirroring pixiv whose `title` is the artwork title
rather than post text.

Also records two API observations in comments/docs: an id that cannot
exist answers `4101105 请求数据发生错误` (kept on the permanent arm), and
the feed endpoints strip `desc.text` so only the detail endpoint shows
whether a post has text.
2026-09-17 22:12:53 +08:00
YoursFunny c1f5d3ca54 feat(sites): add bilibili dynamic support (images and animated images)
Fetch `t.bilibili.com/<id>`, `www.bilibili.com/opus/<id>`,
`t.bilibili.com/h5/dynamic/detail/<id>` and `m.bilibili.com/dynamic/<id>`
through the anonymous `/x/polymer/web-dynamic/v1/detail` endpoint (no
cookie, no WBI signature; the site adds the device cookies
`/x/frontend/finger/spi` hands out, which is what lifts bilibili's
`-352` risk control).

Media: the `major.draw` grid (`.gif` sources become animations, the rest
photos with a downscaled `@518w.jpg` thumbnail used both as preview and
as the oversized fallback), an attached video's cover, and the quoted
dynamic's media for forwards. The video stream itself is not resolved;
`b23.tv` short links stay unmatched (they mostly point at videos, so
matching them would turn a silently ignored link into a failure reply).

`-352`/`-412` map to a retryable error so the queue backs off instead of
dropping the post; a removed dynamic (`500`) is permanent.

Registry-driven, so no bot-side code changes beyond the site lists in the
command replies; found while researching nazurin and
telegram-bili-feed-helper (see BILIBILI_PLAN.md).
2026-09-17 20:48:15 +08:00
YoursFunny 1bb6968108 Merge pull request #3 from TheFunny/dependabot/cargo/rand-0.10.2
build(deps): bump rand from 0.8.8 to 0.10.2
2026-09-17 17:05:30 +08:00
YoursFunny 14b444d109 Merge master into the rand 0.10 migration
Resolve the manifest and lockfile overlap with the rusqlite 0.40 and zip 8.6
bumps that landed on master after this branch was opened:
- crates/xmedia-bot/Cargo.toml: keep rusqlite 0.40 and rand 0.10
- Cargo.lock: re-point x-media/xmedia-bot at rand 0.10.2; teloxide keeps its
  own rand 0.8.8 (it requires ^0.8.5), and cargo pruned the now-unused
  version_check entry

Verified on the merged tree: cargo metadata --locked, cargo fmt --check,
cargo clippy --workspace --all-targets --locked -- -D warnings,
cargo test --workspace --locked (69 + 73 tests).
2026-09-17 16:59:27 +08:00
YoursFunny de3105d4cd Merge pull request #2 from TheFunny/dependabot/cargo/zip-8.6.0
build(deps): bump zip from 2.4.2 to 8.6.0
2026-09-17 16:54:44 +08:00
YoursFunny c46103a23f Merge pull request #1 from TheFunny/dependabot/cargo/rusqlite-0.40.2
build(deps): bump rusqlite from 0.32.1 to 0.40.2
2026-09-17 16:54:00 +08:00
YoursFunny 9529f64b41 fix(send): migrate rand usage to the 0.10 API
rand 0.10 renamed the entry points dependabot's version bump alone cannot
follow: `thread_rng()` -> `rng()`, `gen_range()` -> `random_range()` and the
`Rng` trait -> `RngExt`.

The jitter only needs one float, so use the free function
`rand::random_range(0.2..0.8)` and drop the now-unused trait import instead
of importing `RngExt`.

Verified: cargo fmt --check, cargo clippy --workspace --all-targets -- -D
warnings, cargo test --workspace --locked (retry_delay_seconds_bounds keeps
the 0.2..0.8 jitter window).
2026-09-17 16:43:01 +08:00
dependabot[bot] d0de17329d build(deps): bump rand from 0.8.8 to 0.10.2
Bumps [rand](https://github.com/rust-random/rand) from 0.8.8 to 0.10.2.
- [Release notes](https://github.com/rust-random/rand/releases)
- [Changelog](https://github.com/rust-random/rand/blob/master/CHANGELOG.md)
- [Commits](https://github.com/rust-random/rand/compare/0.8.8...0.10.2)

---
updated-dependencies:
- dependency-name: rand
  dependency-version: 0.10.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-17 03:43:17 +00:00
dependabot[bot] 9af37e92b4 build(deps): bump zip from 2.4.2 to 8.6.0
Bumps [zip](https://github.com/zip-rs/zip2) from 2.4.2 to 8.6.0.
- [Release notes](https://github.com/zip-rs/zip2/releases)
- [Changelog](https://github.com/zip-rs/zip2/blob/master/CHANGELOG.md)
- [Commits](https://github.com/zip-rs/zip2/compare/v2.4.2...v8.6.0)

---
updated-dependencies:
- dependency-name: zip
  dependency-version: 8.6.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-17 03:43:10 +00:00
dependabot[bot] e68a1dbd30 build(deps): bump rusqlite from 0.32.1 to 0.40.2
Bumps [rusqlite](https://github.com/rusqlite/rusqlite) from 0.32.1 to 0.40.2.
- [Release notes](https://github.com/rusqlite/rusqlite/releases)
- [Changelog](https://github.com/rusqlite/rusqlite/blob/master/Changelog.md)
- [Commits](https://github.com/rusqlite/rusqlite/compare/v0.32.1...v0.40.2)

---
updated-dependencies:
- dependency-name: rusqlite
  dependency-version: 0.40.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-17 03:43:05 +00:00
YoursFunny b9c6d16ff0 docs(ci): note why the docker job skips actions/checkout
build-push-action defaults to the Git context, so BuildKit clones the
repo itself and the job never needs the workspace.
2026-09-17 11:38:49 +08:00
32 changed files with 3345 additions and 441 deletions
+5
View File
@@ -80,6 +80,11 @@ jobs:
echo "build=true" >> "$GITHUB_OUTPUT"
fi
# No `actions/checkout` here on purpose: `docker/build-push-action` defaults
# to the Git context (`https://github.com/<owner>/<repo>.git#<ref>`), so
# BuildKit clones the repo itself and authenticates with the automatic
# github.token. Adding `context: .` below without a checkout step would hand
# BuildKit an empty workspace.
docker:
needs: should-build
if: needs.should-build.outputs.build == 'true'
+22 -16
View File
@@ -2,9 +2,9 @@
## Project Overview
Telegram bot (teloxide) that turns post links from X/Twitter, Pixiv, Bluesky, and Misskey (misskey.io) into media messages (images, video, GIF) with the post's title, author, and tags. It supports batch media splitting, retry with persistence, inline queries, forward-channel rebinding with caption templates, and Pixiv ugoira→MP4 transcoding. README is in Chinese; user-facing bot strings are in English. The project is a Rust port of a Python predecessor (see `queue.rs` comments referencing `utils/task_queue.py`).
Telegram bot (teloxide) that turns post links from X/Twitter, Pixiv, Bluesky, Misskey (misskey.io), and Bilibili dynamics into media messages (images, video, GIF) with the post's title, author, and tags. It supports batch media splitting, retry with persistence, inline queries, forward-channel rebinding with caption templates, and Pixiv ugoira→MP4 transcoding. README is in Chinese; user-facing bot strings are in English. The project is a Rust port of a Python predecessor (see `queue.rs` comments referencing `utils/task_queue.py`).
Two-crate Cargo workspace (both v1.6.0, edition 2024, resolver 3):
Two-crate Cargo workspace (both v1.8.0, edition 2024, resolver 3):
- **`crates/x-media`** — library that fetches and normalizes media from the four sites. Pure, no Telegram knowledge.
- **`crates/xmedia-bot`** — the bot binary: teloxide dispatcher, SQLite-backed chat state, persistent task queue.
@@ -18,24 +18,30 @@ Telegram update → Dispatcher (polling or axum webhook) → dptree branches
└─ callback_query → "forward" (copy to channel) / "template|<name>" (apply caption template)
```
Message flow: `message_handler` extracts URLs (from `url`/`text_link` entities, text + caption, deduped) → `x_media::site::fetch(url)``Fetched` → builds a `Task``send::send_media_sequence` (media groups ≤ 9, caption on first item) or `send::send_animation`. On Telegram URL-fetch failure or size error (`send_batch_via_upload`): download via `x_media::site::download_media` to a temp file (≤ 10 MiB), sniff magic bytes (`sniff_ext`), upload via multipart; oversized items fall back to `fallback_url`. On failure: `enqueue_retry` persists resume-state `Task` into the SQLite queue → workers lease (120 s lock TTL) → retry with exponential backoff (≤ 30 s, `MAX_RETRIES = 2`) → dead-letter → `notify_failure`. Success → `post_send_actions`: edit-before-forward prompt with inline buttons, or `copy_messages` to the bound forward channel.
Message flow: `message_handler` extracts URLs (from `url`/`text_link` entities, text + caption, deduped) → `x_media::site::fetch(url)``Fetched` → builds a `Task``send::send_media_sequence` (media groups ≤ 10, caption on first item) or `send::send_animation`. On Telegram URL-fetch failure or size error (`send_batch_via_upload`): download via `x_media::site::download_media` to a temp file (≤ 10 MiB), sniff magic bytes (`sniff_ext`), upload via multipart; oversized items fall back to `fallback_url`. On failure: `enqueue_retry` persists resume-state `Task` into the SQLite queue → workers lease (120 s lock TTL) → retry with exponential backoff (≤ 30 s, `MAX_RETRIES = 2`) → dead-letter → `notify_failure`. Success → `post_send_actions`: edit-before-forward prompt with inline buttons, or `copy_messages` to the bound forward channel.
Debug command: `/debug <url>` runs the same `x_media::site::fetch` and replies with `debug_report` (`handlers/commands.rs`) — site id, normalized cache key, source URL, title/author/tags, sensitive flag, caption and the media list — nothing is sent, cached or forwarded; the report is capped at 4000 chars and sent with HTML parse mode: raw fields are escaped, and the caption is wrapped in a `<blockquote>` so it renders exactly like the sent media caption (escaped text and links included).
Debug command: `/debug <url>` runs the same `x_media::site::fetch` and replies with `debug_report` (`handlers/commands.rs`) — site id, normalized cache key, source URL, title/author/tags, sensitive flag, caption and the media list — nothing is sent, cached or forwarded; the report is capped at 4000 chars and sent with HTML parse mode: raw fields are escaped, and the caption is wrapped in a `<blockquote>` so it renders exactly like the sent media caption (escaped text and links included). The caption it shows is `preview_caption`'s: the chat's per-site format override plus the long-post quoting, i.e. exactly what the send paths produce — showing the raw built-in caption made `/set_format` look like a no-op, and the `/set_format` success reply points users at `/debug` to preview.
The `/test <url>` command runs the ordinary link pipeline (`urls::url_media`) with `PostSend::Suppressed`: the media is sent and cached like any other link, but the chat's `forward_channel_id`/`edit_before_forward` are ignored, so a test never forwards to the channel and never opens the edit prompt (retries and dead-letter notifications behave as usual). Both commands use a custom `parse_arg_remainder` parser (whole remainder, trimmed) because teloxide's built-in `split` parser takes exactly one space-separated token.
User-facing failure text is a function of the error class, never one generic sentence: `urls::fetch_error_message` maps `FetchError::NotFound` (post gone), `Sensitive` (withheld, needs `TWITTER_AUTH_TOKEN`), `Blocked` (source risk control), `Disabled { site }` (a registered site switched off — pixiv without a token, the one case `fetch` answers `Err` instead of `Ok(None)`) and `Transient`/`Http` (source down) apart. The same distinction drives the group hint: a supported link posted in a group (not a channel) gets one `GROUP_LINK_HINT` reply, because the link pipeline is private-chat only.
The `x-media` library: `site::fetch(url)` dispatches through the `SITES` registry (per-site `impl Site`, in order twitter → bsky → misskey → pixiv) and returns `Ok(None)` for unmatched URLs. `Fetched { source_url, caption, title, media: Vec<Media>, sensitive, site_id, … }`; `caption_with(format)` substitutes `{url} {author} {author_url} {title} {tags}`.
The `/test <url>` command runs the ordinary link pipeline (`urls::url_media`) with `PostSend::Suppressed`: the media is sent and cached like any other link, but the chat's `forward_channel_id`/`edit_before_forward` are ignored, so a test never forwards to the channel and never opens the edit prompt (retries and dead-letter notifications behave as usual). `/test`, `/debug`, `/set_format` and `/clear_cache` use the custom `parse_arg_remainder` parser (whole remainder, trimmed) because teloxide's built-in `split` parser takes exactly one space-separated token per field: `/set_format <site> <format>` never parsed with it (and `/clear_cache` without an argument did not either), and a command that fails to parse falls through to the URL flow in silence. `commands::tests::every_documented_invocation_parses` pins every documented form against exactly that.
The inline path (`handlers/inline.rs`) hands media URLs straight to Telegram, which fetches them itself and cannot send site-specific headers — so `x_media::site::needs_media_headers(url)` (true exactly where a site's `media_headers` is non-empty, i.e. pixiv's pximg.net) marks the media that must be skipped instead of shipped broken; locally produced media (ugoira MP4, bsky remux) fails `Url::parse` and is skipped the same way. Inline results are therefore URL-only by construction.
`url_media` is a thin wrapper over `url_media_inner`: `run_with_chat_action` sends the chat action, then re-sends it every `ACTION_REFRESH` (4 s) while the pipeline future is pending, because Telegram drops an action after ~5 s and a fetch (ugoira encode, HLS remux) plus an upload routinely outlasts that. The pipeline flips the shared `ActionHint` from `Typing` to `UploadPhoto`/`UploadVideo` once the media kinds are known. The `select!` is `biased` on the pipeline branch so a finished pipeline never emits a stray action.
The `x-media` library: `site::fetch(url)` dispatches through the `SITES` registry (per-site `impl Site`, in order twitter → bsky → misskey → pixiv → bilibili) and returns `Ok(None)` for unmatched URLs (`Err(FetchError::Disabled { site })` when the URL matches a registered site whose `enabled()` is false — see `disabled_site`). `Fetched { source_url, caption, title, content, media: Vec<Media>, sensitive, site_id, … }` (title and content are split per platform: a pixiv artwork's title and description, a bilibili headline and body, and text-only posts whose text is all `content`); `caption_with(format)` substitutes `{url} {author} {author_url} {title} {content} {tags}`.
## Key Directories
| Path | Purpose |
|---|---|
| `crates/x-media/src/` | Fetch library. `site/mod.rs` = dispatcher + `Fetched`/`FetchError`/`download_media`/`media_size`; `media.rs` = `Media` enum; `examples/fetch.rs` = end-to-end usage sample |
| `crates/x-media/src/site/<twitter\|pixiv\|bsky\|misskey>/` | One directory per site: `mod.rs` (re-exports), `interface.rs` (PATTERN, `enabled()`, `fetch_from_url()`, `cache_key`/`is_retryable`/`media_headers`, unit struct `<Name>Site` implementing `site::Site`, `From<SiteStruct> for Fetched`), `model.rs` (serde DTOs). Pixiv adds `api.rs` (auth + transport); twitter adds `auth.rs` (logged-in GraphQL `TweetDetail` fallback for NSFW tweets, gated on `TWITTER_AUTH_TOKEN`). Misskey targets misskey.io only (`POST /api/notes/show`, 400+`NO_SUCH_NOTE` → NotFound). Twitter's `from_syndication_json` HTML-decodes the API text — syndication and GraphQL `full_text` both arrive pre-escaped (`&gt;` `&lt;` `&amp;` `&#39;`) — so the stored text is raw and the caption escapes exactly once |
| `crates/xmedia-bot/src/main.rs` | Entry point: env/log init, command registration (`register_commands`), shared `send::BOT` force-init, queue worker start, site login validation (`site::validate_all`), 300 s edit-expiry sweep, dptree handler tree, webhook vs polling dispatch |
| `crates/x-media/src/site/<twitter\|pixiv\|bsky\|misskey\|bilibili>/` | One directory per site: `mod.rs` (re-exports), `interface.rs` (PATTERN, `enabled()`, `fetch_from_url()`, `cache_key`/`is_retryable`/`media_headers`, unit struct `<Name>Site` implementing `site::Site`, `From<SiteStruct> for Fetched`), `model.rs` (serde DTOs). Pixiv adds `api.rs` (auth + transport); twitter adds `auth.rs` (logged-in GraphQL `TweetDetail` fallback for NSFW tweets, gated on `TWITTER_AUTH_TOKEN`; without the token a withheld tweet stays `FetchError::Sensitive` and the bot reports it as age-restricted instead of "no media"). Misskey targets misskey.io only (`POST /api/notes/show`, 400+`NO_SUCH_NOTE` → NotFound). Bilibili fetches dynamics (images/animated images only — an attached video degrades to its cover, and its title stands in for the post text, which AV dynamics do not have) from `/x/polymer/web-dynamic/v1/detail` sent with `features=itemOpusStyle` (without that flag the legacy serialization drops an image/text post's body and headline entirely — `desc` comes back `null`; the adapter still parses the legacy `major.draw`/`desc`/`archive` shapes as a fallback). No WBI signature is involved; device cookies `buvid3`/`buvid4` are fetched automatically from `/x/frontend/finger/spi` because bilibili's `-352` risk control starts rejecting plain requests, `BILIBILI_COOKIE` is the escalation when an IP stays blocked; `b23.tv` short links are deliberately unmatched. Twitter's `from_syndication_json` HTML-decodes the API text — syndication and GraphQL `full_text` both arrive pre-escaped (`&gt;` `&lt;` `&amp;` `&#39;`) — so the stored text is raw and the caption escap
| `crates/xmedia-bot/src/main.rs` | Entry point: env/log init, command registration (`register_commands``setMyCommands` plus the profile description texts), shared `send::BOT` force-init, queue worker start, site login validation (`site::validate_all`), 300 s edit-expiry sweep (expired prompts are rewritten in place to `EDIT_PROMPT_EXPIRED_TEXT` with an empty keyboard — an edit, never a new message, so a background timer cannot wake a chat), dptree handler tree, webhook vs polling dispatch |
| `crates/xmedia-bot/src/config.rs` | Manual env parsing into `Config` |
| `crates/xmedia-bot/src/db.rs` | `DbPool`: one shared SQLite connection pool (`POOL_SIZE = 4`, WAL, busy_timeout) for all three tables over `$DATA_DIR/task_queue.db` (default `data/`) — the three stores share it; `open_store` creates file + schema, `with_conn` runs all rusqlite I/O in `spawn_blocking` |
| `crates/xmedia-bot/src/handlers/` | Handler modules: `mod.rs` (message entry point, `reply`, `log_key`), `commands.rs` (teloxide `BotCommands` enum + command executor, incl. `/test <url>` (send-only) / `/debug <url>` (parse-only) and the admin-only `/bot_dict` state dump), `urls.rs` (URL extraction + bounded job channel (256) drained by `URL_WORKERS = 8` workers (`start_url_workers`) — backpressure instead of unbounded spawns; teloxide's per-chat workers are sequential — batch-forwards need concurrency), `inline.rs`/`callback.rs` (inline queries / edit-before-forward buttons), `statics.rs` (global statics) |
| `crates/xmedia-bot/src/handlers/` | Handler modules: `mod.rs` (message entry point, `reply`, `log_key`, the group-only `GROUP_LINK_HINT` for a supported link posted outside a private chat), `commands.rs` (teloxide `BotCommands` enum + command executor, incl. `/test <url>` (send-only) / `/debug <url>` (parse-only) and the admin-only `/bot_dict` state dump; `/set_format` rejects unknown `{…}` placeholders and resets with `-`), `urls.rs` (URL extraction + bounded job channel (256) drained by `URL_WORKERS = 8` workers (`start_url_workers`) — backpressure instead of unbounded spawns; teloxide's per-chat workers are sequential — batch-forwards need concurrency), `inline.rs`/`callback.rs` (inline queries / edit-before-forward buttons, incl. `skip`), `statics.rs` (global statics) |
| `crates/xmedia-bot/src/state.rs` | `ChatStore`: parking_lot `Mutex<HashMap>` cache + SQLite write-through (`chat_state` table) |
| `crates/xmedia-bot/src/link_cache.rs` | `LinkCache`: SQLite-backed cache (`link_cache` table) of successfully sent posts — raw caption fields + Telegram `file_id`s; repeat links re-send locally (no fetch/upload), TTL + prune, invalidated on permanent send failure |
| `crates/xmedia-bot/src/queue.rs` | `PersistentTaskQueue`: SQLite-backed queue (`tasks` table), `QUEUE_WORKERS = 4` concurrent workers (lease via `BEGIN IMMEDIATE` + `locked_until` TTL), retry→dead-letter, `notify_one` worker wakeup plus a separate `Notify` for the 30 s lease-expiry sweep (a shared one let the sweep steal the workers' wakeup permit), `busy_timeout` on all connections |
@@ -60,7 +66,7 @@ Docker: `docker build -t tgxmb .` then `docker run --rm -d --name tgxmb --env-fi
## Code Conventions & Common Patterns
- **Errors via `thiserror` derive** (no anyhow): the public, stringified errors — `FetchError` (`Http`/`Json`/`Pixiv`/`Site`/`NotFound`/`Blocked`) and `PixivError` — derive `thiserror::Error` with `#[from]` conversions; `Display`/`source()` come from the derive. The internal control-flow enums — `QueueError` (`Retryable { delay_seconds, payload }` / `Permanent`), `SendError` (Retryable/Permanent), `Classification`, `FallbackError` — carry no `Display` and are handled by direct variant matching. New errors should follow the same split: stringified/public errors derive `thiserror`, internal flow enums stay plain.
- **Errors via `thiserror` derive** (no anyhow): the public, stringified errors — `FetchError` (`Http`/`Json`/`Pixiv`/`Site`/`NotFound`/`Blocked`/`Disabled`/`Sensitive`/`TooLarge`/`Transient`/`Io`) and `PixivError` — derive `thiserror::Error` with `#[from]` conversions; `Display`/`source()` come from the derive. The internal control-flow enums — `QueueError` (`Retryable { delay_seconds, payload }` / `Permanent`), `SendError` (Retryable/Permanent), `Classification`, `FallbackError` — carry no `Display` and are handled by direct variant matching. New errors should follow the same split: stringified/public errors derive `thiserror`, internal flow enums stay plain.
- **Global state via `std::sync::LazyLock` statics**, not DI: `CONFIG`, `CHAT_STORE`, `TASK_QUEUE` in `handlers/statics.rs`; shared reqwest `CLIENT` in `x-media/src/site/mod.rs`. `Bot` is passed/cloned into handlers; queue workers share the process-wide `send::BOT` (`LazyLock<Bot>`, force-initialized in `main` so a missing token fails at startup).
- **Async**: tokio multi-thread runtime (`#[tokio::main]` default). All rusqlite I/O inside `tokio::task::spawn_blocking`. Long loops use `tokio::select!` with `tokio::sync::{watch, Notify}` stop/wake channels. No streams.
- **Blocking sync primitives**: `parking_lot::Mutex` for hot caches, `tokio::sync::Mutex` for async-shared state (pixiv token cache), `AtomicBool` for feature gates.
@@ -75,10 +81,10 @@ Docker: `docker build -t tgxmb .` then `docker run --rm -d --name tgxmb --env-fi
| File | Why it matters |
|---|---|
| `crates/xmedia-bot/src/main.rs` | Startup sequence, webhook vs polling, graceful shutdown (SIGINT via teloxide ctrlc / SIGTERM via `stop_token` for docker, → sweep stop → admin msg → queue stop) |
| `crates/xmedia-bot/src/handlers/` | `statics.rs` = `CHAT_STORE`/`TASK_QUEUE`/`CONFIG` singletons (open `$DATA_DIR/task_queue.db`, default `data/` **relative to CWD**, dir auto-created); `commands.rs` = command dispatch (incl. `/test <url>` send-only, `/debug <url>` parse-only, and the admin-only `/bot_dict` state dump); `urls.rs` = URL extraction + the per-URL pipeline (`url_media` takes a `PostSend` mode: chat settings vs `/test`'s suppressed actions); `inline.rs` = debounced inline queries; `callback.rs` = edit-before-forward buttons (dptree entry + testable `handle_callback` core) |
| `crates/xmedia-bot/src/send/` | `mod.rs`: constants `MAX_MEDIA_GROUP = 9`; `classify_request_error`; the senders. `upload.rs`: download-and-reupload fallback triggered only by Telegram API errors (`is_media_fetch_failure` / `is_size_error`). `post_send.rs`: settlement (`settle_task`), cache write, post-send actions, queue handlers. `input_media.rs`: payload → `InputMedia` |
| `crates/xmedia-bot/src/handlers/` | `statics.rs` = `CHAT_STORE`/`TASK_QUEUE`/`CONFIG` singletons (open `$DATA_DIR/task_queue.db`, default `data/` **relative to CWD**, dir auto-created); `commands.rs` = command dispatch (incl. `/test <url>` send-only, `/debug <url>` parse-only, the read-only `/settings` every chat member can read — unlike the admin-only `/bot_dict` raw dump — and template removal; `/start`/`/help` carry the guidance teloxide's `descriptions()` cannot render, and `/set_format` rejects unknown `{…}` placeholders, resetting with `-`); `urls.rs` = URL extraction + the per-URL pipeline (`url_media` takes a `PostSend` mode: chat settings vs `/test`'s suppressed actions); `inline.rs` = debounced inline queries (hotlink-protected and local media skipped); `callback.rs` = edit-before-forward buttons (dptree entry + testable `handle_callback` core, incl. `skip`) |
| `crates/xmedia-bot/src/send/` | `mod.rs`: constants `MAX_MEDIA_GROUP = 10`; `classify_request_error`; the senders. `upload.rs`: download-and-reupload fallback triggered only by Telegram API errors (`is_media_fetch_failure` / `is_size_error`). `post_send.rs`: settlement (`settle_task`), cache write, post-send actions (dead-letter text via `failure_text`: post key + cause, since the raw error alone does not say which link died), queue handlers. `input_media.rs`: payload → `InputMedia` |
| `crates/xmedia-bot/src/photo.rs` | Pure-Rust photo processing (no ffmpeg): `png` (image-png) decode/encode + `zune-jpeg` decode + `fast_image_resize` Lanczos3 downscale + `jpeg-encoder`. Photos over Telegram's limits (width + height > 10000 px → `PHOTO_INVALID_DIMENSIONS`; bytes > 10 MiB) are decoded, downscaled keeping the format, PNG bit depth > 24 (RGBA 32-bit / 16-bit per channel) reduced to 24-bit RGB with alpha flattened white (≤24-bit untouched, never upconverted), and transcoded to JPEG only if still over the cap; memory budget guarded, otherwise the item's smaller fallback URL |
| `crates/x-media/src/site/mod.rs` | Dispatcher, `Fetched`/`FetchError`, shared `CLIENT`, `download_media` (adds `Referer: https://www.pixiv.net/` for `pximg.net` hotlink protection) |
| `crates/x-media/src/site/mod.rs` | Dispatcher, `Fetched`/`FetchError`, shared `CLIENT`, `download_media` (adds `Referer: https://www.pixiv.net/` for `pximg.net` hotlink protection), `needs_media_headers` (the same per-site rule, asked by the inline path to skip what Telegram cannot fetch) |
| `crates/x-media/src/site/pixiv/api.rs` | OAuth token exchange (hardcoded app client id/secret), access-token cache, ugoira zip→MP4 via ffmpeg in `spawn_blocking` |
| `Dockerfile` | Multi-stage: cached dep layer via stub sources + `touch *.rs` mtime bump (cargo's freshness is mtime-based and `cargo clean -p` removes 0 files — the touch is what forces the real sources to rebuild while deps stay cached), static ffmpeg from ffmpeg.martin-riedl.de (`FFMPEG_URL` arg, optional `FFMPEG_SHA256` checksum, `unzip -t` integrity check), `debian:bookworm-slim` runtime, entrypoint. Runtime ships **no libssl/libcrypto/CA bundle** — rustls webpki-roots handles all TLS, and the static ffmpeg only processes local files (downloads go through reqwest) |
| `docker-entrypoint.sh` | Privilege drop: `useradd` with `LOCAL_USER_ID` (default 9001) + `setpriv` (no gosu on bookworm-slim) |
@@ -92,16 +98,16 @@ Docker: `docker build -t tgxmb .` then `docker run --rm -d --name tgxmb --env-fi
- Package manager: **Cargo** (workspace with path dep `x-media``xmedia-bot`). No `[workspace.package]`/shared deps — each crate lists deps independently.
- **TLS is rustls end-to-end** (no native-tls/openssl in the tree, no libssl in the Docker runtime image): `teloxide` is declared `default-features = false` with `["webhooks-axum", "macros", "rustls", "ctrlc_handler"]` (the removed `default` also carried `native-tls` and `ctrlc_handler` — the latter must stay); x-media's reqwest is `default-features = false` with `["json", "rustls-tls"]` (webpki-roots baked in, so the image ships no CA bundle). One reqwest 0.12.28 in the lock.
- **Versioning**: bump the version in all three places (`crates/x-media/Cargo.toml`, `crates/xmedia-bot/Cargo.toml`, `Cargo.lock`) and **keep `README.md`, `README.en.md` and `AGENTS.md` in sync with the code on every bump**, then commit (`chore: bump version to X.Y.Z`), create an annotated tag `vX.Y.Z`, and push branch + tag (the tag push triggers the Docker Hub build). The tag must equal both crate versions: `.github/workflows/docker.yml` verifies that before building, and `--locked` verifies the lock file.
- Config is **environment-variable driven** (dotenv loads `.env`, gitignored; no `.env.example` exists). Key vars: `TELOXIDE_TOKEN` (required), `PIXIV_REFRESH_TOKEN`, `TWITTER_AUTH_TOKEN` (optional; x.com `auth_token` cookie — enables the logged-in GraphQL fallback that fetches NSFW tweets syndication withholds), `BOT_ADMIN` (comma-separated ids), `EDIT_MESSAGE_TTL_SECONDS` (default 86400), `LINK_CACHE_TTL_SECONDS` (default 604800), `DATA_DIR` (default `data`, CWD-relative; the SQLite dir, auto-created), `WEBHOOK`/`WEBHOOK_URL`/`WEBHOOK_LISTEN`/`WEBHOOK_PORT`/`WEBHOOK_CERT`/`WEBHOOK_SECRET_TOKEN` (webhook mode requires URL/listen/port, `.expect`ed; `WEBHOOK_CERT` is Telegram-facing self-signed validation only — TLS must be terminated by a reverse proxy), `RUST_LOG`, `TELOXIDE_PROXY`, `LOCAL_USER_ID` (entrypoint only).
- Config is **environment-variable driven** (dotenv loads `.env`, gitignored; no `.env.example` exists). Key vars: `TELOXIDE_TOKEN` (required), `PIXIV_REFRESH_TOKEN`, `TWITTER_AUTH_TOKEN` (optional; x.com `auth_token` cookie — enables the logged-in GraphQL fallback that fetches NSFW tweets syndication withholds), `BILIBILI_COOKIE` (optional; whole bilibili cookie string — bilibili dynamics fetch anonymously and add their own device cookies, this only rescues an egress IP that bilibili has hard-flagged with `-352`/412), `BOT_ADMIN` (comma-separated ids), `EDIT_MESSAGE_TTL_SECONDS` (default 86400), `LINK_CACHE_TTL_SECONDS` (default 604800), `CAPTION_QUOTE_TEXT_CHARS` (default 200; a post whose text — the `title` plus `content` joined, see `site::compose_text` — reaches this length gets that text wrapped in an expandable blockquote inside its caption, the URL and author line staying outside; `0` disables it. Applied at the send boundary in `send::quote_long_caption`, which locates the text as what follows the author link, so a `/set_format` that moves `{title}`/`{content}` elsewhere and pixiv's title-inside-a-link layout opt out; `copy_messages` forwards and queued retries inherit the wrap, while the edit-before-forward rewrite stays unquoted by design), `DATA_DIR` (default `data`, CWD-relative; the SQLite dir, auto-created), `WEBHOOK`/`WEBHOOK_URL`/`WEBHOOK_LISTEN`/`WEBHOOK_PORT`/`WEBHOOK_CERT`/`WEBHOOK_SECRET_TOKEN` (webhook mode requires URL/listen/port, `.expect`ed; `WEBHOOK_CERT` is Telegram-facing self-signed validation only — TLS must be terminated by a reverse proxy), `RUST_LOG`, `TELOXIDE_PROXY`, `LOCAL_USER_ID` (entrypoint only).
- SQLite via `rusqlite` with `bundled` feature (no system libsqlite needed). DB file `$DATA_DIR/task_queue.db` (default `data/task_queue.db`, CWD-relative — run from the workspace root, or `/app` in Docker; set `DATA_DIR` to pin state anywhere). Mount `./data` and `./cert` volumes.
- `.gitattributes` enforces LF for `*.sh` (CRLF breaks shebangs in containers). `.gitignore`: `.env`, `data/`, `cert/`, `docker-compose.yml`, `/target`, `.idea/`.
- Docs are in Chinese (README, AGENTS.md); user-facing bot strings are in English. Keep that split when editing user-facing strings and docs.
## Testing & QA
- **~135 tests, all inline `#[cfg(test)] mod tests`** — no `tests/` integration directories. Framework: built-in Rust test + `#[tokio::test]` (dev-deps only in `x-media`: tokio macros/rt-multi-thread, dotenv).
- **~200 tests, all inline `#[cfg(test)] mod tests`** — no `tests/` integration directories. Framework: built-in Rust test + `#[tokio::test]` (dev-deps only in `x-media`: tokio macros/rt-multi-thread, dotenv).
- No mocking framework anywhere (no mockito/wiremock/mockall). Conventions: pure-function units (regex parsing, serde round-trips, chunking, retry math) tested synchronously; async tests use real dependencies — file-backed SQLite via `tempfile` (`queue.rs::new_queue()` helper), live network fetches.
- Live-network tests exist in `site/twitter/interface.rs` (5), `site/bsky/interface.rs` (2), `site/misskey/interface.rs` (1), `site/pixiv/api.rs` (1); `photo.rs` adds one `#[ignore = "heavy: …"]` test. `site/mod.rs` also has a **token-gated but not `#[ignore]`d** pixiv download test (`download_media_pixiv_original_with_referer`): it hits `i.pximg.net` whenever `PIXIV_REFRESH_TOKEN` is set, so a local `cargo test --workspace` is not fully offline and can flake on a pixiv CDN body timeout. Test gating convention (enforced by `.github/workflows/ci.yml`): pure unit tests always run; live-network tests carry `#[ignore = "live network: ..."]` (run via `cargo test --workspace -- --ignored live`); token-gated pixiv tests early-return when `PIXIV_REFRESH_TOKEN` is absent **or empty** (an unset GitHub secret arrives as `""``is_err()` alone would run them tokenless and fail). Run the full offline suite with `cargo test --workspace`.
- Live-network tests exist in `site/twitter/interface.rs` (5), `site/bsky/interface.rs` (2), `site/misskey/interface.rs` (1), `site/bilibili/interface.rs` (4), `site/pixiv/api.rs` (1); `photo.rs` adds one `#[ignore = "heavy: …"]` test. `site/mod.rs` also has a **token-gated but not `#[ignore]`d** pixiv download test (`download_media_pixiv_original_with_referer`): it hits `i.pximg.net` whenever `PIXIV_REFRESH_TOKEN` is set, so a local `cargo test --workspace` is not fully offline and can flake on a pixiv CDN body timeout. `disabled_site_is_reported_not_ignored` (same file) is gated the other way round: it asserts `fetch` answers `FetchError::Disabled { site: "pixiv" }` for a pixiv link and early-returns when `PIXIV_REFRESH_TOKEN` **is** set (the site is then enabled). Test gating convention (enforced by `.github/workflows/ci.yml`): pure unit tests always run; live-network tests carry `#[ignore = "live network: ..."]` (run via `cargo test --workspace -- --ignored live`); token-gated pixiv tests early-return when `PIXIV_REFRESH_TOKEN` is absent **or empty** (an unset GitHub secret arrives as `""``is_err()` alone would run them tokenless and fail), and the bilibili live tests early-return when the API answers risk control (`-352`, which bilibili applies per IP by request volume). Run the full offline suite with `cargo test --workspace`.
- Fixtures are inline `serde_json::json!` builder fns (`fixture()`, `thread_json()`, `illust_json()`), not files. The shared `CLIENT` sets `pool_max_idle_per_host(0)` under `#[cfg(test)]` to avoid cross-runtime `DispatchGone`.
- **CI** — `.github/workflows/ci.yml` (actions pinned to commit SHAs, `--locked` on every cargo invocation, `concurrency` cancels superseded runs, `RUST_BACKTRACE=1`) runs `cargo fmt --check` + `cargo clippy --workspace --all-targets --locked -- -D warnings` + `cargo test --workspace --locked` + a release-profile `cargo build --release --locked` + an `actions-rust-lang/audit` dependency-vulnerability gate (offline, no secrets, on every push/PR) and a `live` job (schedule/manual/tag only, `-p x-media` since every network/secret-gated test lives there, `continue-on-error`) for the `#[ignore]`d live + token tests. `.github/workflows/docker.yml` builds and pushes the image on master/tag and runs a **build-only check on pull requests touching the build inputs** (`Dockerfile`, entrypoint, manifests, `.dockerignore`); a release tag must match both crate versions or the build stops, and `FFMPEG_URL`/`FFMPEG_SHA256` are taken from repository variables when set (a release can pin an exact ffmpeg build). `.github/dependabot.yml` keeps crates, the pinned actions and the Docker base images current.
- Untested and hard to test without a mock seam: `main.rs`, `config.rs`, `db.rs`, `handlers/statics.rs`, `media_sender.rs` (holds the `MockSender` itself); in `x-media`: `media.rs`, `lib.rs`, all `model.rs`. The `commands.rs` *executor* needs a real `Bot` (only its pure report builder is tested). Everything else — `handlers/{mod,callback,inline,urls}.rs`, `send/*`, `ctx.rs`, `state.rs`, `queue.rs`, `link_cache.rs`, `rate_limit.rs` — is driven through `TestStores`/`ctx::test_support` and the scripted `MockSender`.
+168
View File
@@ -0,0 +1,168 @@
# Bilibili 动态支持:研究与实现记录
状态:已实现(`crates/x-media/src/site/bilibili/`)。本文记录上游调研、实测数据与最终设计;
长期契约以 `AGENTS.md` 为准。
范围:**只发动态里的图片与动图**。动态内嵌视频不发流,降级为封面图;`b23.tv` 短链不匹配;
视频页 / 番剧 / 直播间 / 专栏 / 音频均不支持。
---
## 1. 上游实现研究
### 1.1 nazurin`nazurin/sites/bilibili/`4 个文件 ~6 KB
- 入口正则:`t\.bilibili\.com/(\d+)``t\.bilibili\.com/h5/dynamic/detail/(\d+)``bilibili\.com/opus/(\d+)`
- 请求:`GET https://api.bilibili.com/x/polymer/web-dynamic/v1/detail?id={id}`,仅加 `Referer: https://t.bilibili.com/{id}`
**无 cookie、无 WBI 签名、无 `build` 参数**
- 错误:`code == 4101147` → not found`code != 0` 或缺 `data` → 报错。
- 媒体:只取 `item.modules.module_dynamic.major.draw.items[].src`;缩略图 `src + "@518w.jpg"`
`size` 字段单位是 **KB**`major` 为空或 `draw.items` 为空 → "No image found"。
**忽略视频、转发(forward)与纯文字动态**
- caption`"#" + module_author.name` + `module_dynamic.desc.text`,链接写死 `https://www.bilibili.com/opus/{id}`
### 1.2 telegram-bili-feed-helper`biliparser/provider/bilibili/`9 个文件 ~57 KB
- 9 个策略类(Video/Opus/Live/Audio/Read + Feed 基类 + Credential + api 工具):门禁正则
`bilibili\.com|b23\.tv|BV\w{10}|av\d+`,再分流,兜底 `client.head(url)` 跟随重定向后按子串分流。
- 动态:`GET /x/polymer/web-dynamic/desktop/v1/detail?id={id}&build=11605`**单条,无分页**);
客户端带桌面 UA、随机 `buvid3={uuid}infoc`;登录态用 `bilibili-api-python``Credential`
Redis 持久化 `SESSDATA/bili_jct/buvid3/buvid4/ac_time_value/DedeUserID`,扫码登录)。
- **同样没有 WBI 签名 / appkey 签名**playurl 用的是非 WBI 的 `/x/player/playurl`
- 媒体:`major.type` 分派 —— DRAW 取全部 `items[].src`ARCHIVE/PGC/ARTICLE/MUSIC/COMMON/LIVE
只取一张 `cover`;FORWARD 取原动态作者/正文并递归进 `orig` 找媒体。
- 视频:仅独立 video 策略解析(`qn` 720P→480P→360P 试 durl,再退 DASH + ffmpeg 合并);
**动态内嵌视频只发封面**
- 错误:要求 `status==200 && code==0`;风控 `-352`/`-412` 无特殊处理。
### 1.3 取舍
| 维度 | nazurin | bff | 本仓库 |
|---|---|---|---|
| 接口 | `v1/detail?id=` | `desktop/v1/detail?id=&build=` | `v1/detail?id=`(实测可用) |
| 认证 | 无 | buvid3 + SESSDATA | 默认匿名;可选 `BILIBILI_COOKIE` |
| WBI | 无 | 无 | 不实现(无需求) |
| 图片 | `major.draw.items` | 同 + forward 递归 | 同,加 `orig` 递归、`http→https``.gif → Animated` |
| 视频 | 完全忽略 | 动态内嵌视频发封面 | 发封面(不发流) |
| 短链 | 不匹配 | 跟随重定向 | 不匹配(多数短链是视频,会让"静默忽略"变成失败提示) |
---
## 2. 实测验证(2026-09-17,真实请求)
| 验证项 | 结果 |
|---|---|
| `v1/detail?id=`(无 cookie、UA `Mozilla/5.0`、带 Referer | `200 {"code":0}` ✅ |
| 同上,不带 cookie 也不带 Referer | `200 {"code":0}` ✅(无强制鉴权) |
| bff 的 `bilibili_pc/…Electron/22.3.27` UA | `code:-352` ❌ → **不要抄它的 UA** |
| `desktop/v1/detail?build=11605` | `code:-352` ❌ |
| `feed/space?host_mid=`(用户时间线) | 首次成功、随后 `-352`,也见过 HTTP 412 → **不碰** |
| 不存在 / 已删除的动态 | `code:500` "Cannot read property 'only_fans' of undefined"nazurin 的 4101147 已失效) |
| 非数字 id | `code:-400` param parsing failed |
| 图片 `i0.hdslb.com/bfs/new_dyn/*.jpg` | `HEAD 200 image/jpeg`,带/不带 Referer 均可;`+@518w.jpg` → 2542 KB ✅ |
| `t.bilibili.com/h5/dynamic/detail/<id>` | `200` ✅ |
| `m.bilibili.com/dynamic/<id>` | `302 → t.bilibili.com/<id>` ✅ |
| `www.bilibili.com/opus/<id>` | `200`,转发动态 `302 → t.bilibili.com/<id>` ✅ |
| `b23.tv/BV1JTtt6JEZu` | `302 → www.bilibili.com/video/BV…`(视频) |
| `b23.tv/<无效码>` | **HTTP 200** + `{"code":-404}` ⚠️ 短链判定不能只看状态码 |
| `playurl`(仅调研用,未采用) | `fnval=1` 匿名给 durl720P=9.18 MiB / 360P=2.97 MiB`fnval=4048` 匿名 DASH 上限仅 480P |
| `dyn_archive` 字段 | 有 `aid/bvid/cover/title/duration_text`**没有 `cid`**(所以发流要再来一次 `view` 请求) |
| **风控阶梯(同一 IP 连续请求后实测)** | ① 无 cookie → `-352`;② 仅 `buvid3` → 仍 `-352`;③ `buvid3`+`buvid4`(取自匿名 `/x/frontend/finger/spi`)→ **`code:0` 恢复**;④ 继续高频请求后 → 连同 buvid 一起 `-352`(此时只有登录 cookie 或换 IP) |
| **正文位置(24 条真实动态逐条审计)** | 有正文的动态都在 `module_dynamic.desc.text`(图文/转发/纯文字,含 34–193 字样本);**AV(视频投稿)动态 `desc` 恒为 `null`**,内容在 `major.archive.title` / `.desc` 卡片里 → 已做 title 回退 |
| **`features=itemOpusStyle` 的效果** | 同一端点带此参数后,图文帖改为 `major.opus` 形态:`pics[]`(图,key 是 `url`)、`summary.text`(正文,未截断,实测 307 字整段)、`title`(可选标题);不带参数则是 legacy `major.draw` + `desc`,而 **opus 图文帖的 `desc` 为 `null`、正文与标题完全丢失**`opus/1248857553488576532`legacy `desc:null`,带参数 `summary.text="[doge_金箍]黑白搭配"`)。AV / 转发帖不受该参数影响 → 适配器改为请求时带参数,并保留 legacy 形态兜底 |
| feed 与 detail 的差异 | `feed/space` 的 item 会把 `desc.text` 挖空,**只有 detail 有正文** → 排查时不要用 feed 数据判断正文缺失 |
| 不存在的 19 位 id | `4101105 请求数据发生错误`(提示可重试,但只出现在不可能存在的 id 上)→ 仍归入永久错误,见 `code_error` 注释 |
测试样本(live 测试用):
| 样本 | id | 期望 |
|---|---|---|
| 图片动态(2 图 + 话题) | `1245284537985925159` | 2 个 `Illustration``{tags}` = `ALin出道20周年快乐` |
| 转发动态 | `1248982077447077907` | 媒体来自 `orig`1 图),正文可含 `//@` |
| 视频动态 | `1248717597691609105` | 封面 1 张 `Illustration` |
| 纯文字动态 | `1246767523595026450` | `media` 为空 |
关键字段路径:
```
data.item.id_str
data.item.modules.module_author.{name,mid}
data.item.modules.module_dynamic.desc.text
data.item.modules.module_dynamic.topic.{id,name} # 单话题,{tags} 来源
data.item.modules.module_dynamic.major.{draw.items[].src, archive.cover}
data.item.orig # 转发时存在,结构与 item 相同
```
---
## 3. 实现
```
crates/x-media/src/site/bilibili/mod.rs # re-export
crates/x-media/src/site/bilibili/interface.rs # PATTERN / cache_key / enabled / is_retryable /
# media_headers / BilibiliSite / fetch / code_error /
# From<Item> for Fetched / caption / 12 单测 + 2 live
crates/x-media/src/site/bilibili/model.rs # 纯 Deserialize DTO(全 Option
```
- **正则**(同时用于分发、抽 id、缓存键,一个正则三用):
`^(?:https?://)?(?:www|t|m)\.bilibili\.com/(?:opus/|dynamic/|h5/dynamic/detail/)?(\d+)`
- **缓存键**`bilibili:<动态 id>``source_url` 统一 `https://www.bilibili.com/opus/{id}`
- **请求**`GET /x/polymer/web-dynamic/v1/detail?id=` + `Referer: https://www.bilibili.com/`
`Cookie` 头按优先级取:`BILIBILI_COOKIE` → 缓存的设备 cookie`GET /x/frontend/finger/spi``buvid3`/`buvid4`
进程内缓存一次;取不到就不带 cookie,仅 debug 日志)→ 无。指纹接口本身失败**不**让抓取失败。
走共享 `CLIENT`UA `Mozilla/5.0`30s 超时,`TELOXIDE_PROXY` 透传)。
- **错误映射**`0` → 成功;`-352/-412` 与 HTTP 412 → `Transient`(可重试,队列退避;首次记一条 warn 提示
`BILIBILI_COOKIE`);`500`/`4101147``NotFound`(永久);其他 code → `Site`(永久)。
- **媒体**
- `major.opus.pics[]`(带 `features=itemOpusStyle` 时的图文帖形态,字段名是 `url`)→ 每张一张图;
其次 `major.draw.items[]`legacy,字段名 `src`)→ 同样逐张;`http://` / `//``https://`,非 https 开头直接丢弃。
`.gif``Media::Animated``thumbnail_url` 留空,Telegram 自己取首帧——`@518w.jpg` 只对 jpg/webp 实测过),
其余 → `Media::Illustration``thumbnail_url = url + "@518w.jpg"`,兼作超大时的降级 URL)。
- `major.archive.cover` → 1 张 `Illustration`(视频不发流)。
- 转发且自身无媒体 → 递归取 `orig` 的媒体;正文拼 `//@{原作者}:\n{原文}`
- 其他 majorPGC/ARTICLE/MUSIC/LIVE/COMMON)不建模 → 无媒体,走既有 "No media found"。
- **正文 / title**(按信息量从多到少回退):`major.opus.title` + `major.opus.summary.text`
`module_dynamic.desc.text``major.archive.title`。三者分别对应:图文文档(标题+正文)、
legacy/转发帖正文、视频投稿卡片标题。开头结尾空白做 trim;整体再由既有 `truncate_caption` 截断。
- **caption**(与 misskey 同形):`{opus 链接}\n<a href="space.bilibili.com/{mid}">{name}</a>: {正文}`
`RenderData``{tags}` 来自话题名;正文由既有 `truncate_caption` 截断。
- **注册表**`SITES` 末尾追加 → `/set_format` 白名单、链接缓存、启动校验、日志前缀全部自动生效。
- **bot 侧仅文案**`handlers/commands.rs` 三处站点清单字符串 + `state.rs`/`handlers/mod.rs` 注释。
### 与原计划的偏差(及原因)
| 原计划 | 实际 | 原因 |
|---|---|---|
| `x/web-interface/view` + `playurl` 发视频 | 不做 | 需求收窄为图片/动图;视频只发封面 |
| `site/mod.rs``MAX_MEDIA_UPLOAD_BYTES` 常量 | 不加 | 没有视频尺寸决策就不需要该常量,避免跨 crate 耦合 |
| `b23.tv` 短链(跟随重定向) | 不匹配 | 多数短链指向视频,匹配后会把"静默忽略"变成用户的 "Failed to fetch media" |
| `validate()` 校验 cookie | 不做 | 匿名可用,cookie 失效不致命;校验要额外请求一个端点,收益低 |
| `media_headers` 给 hdslb 加 Referer | 返回 `None` | 实测图片与 durl 均无需 Referer(注释里记了这条验证) |
| 计划阶段认为设备 cookie 是 YAGNI,不实现 | **实现**`buvid3`+`buvid4`) | 计划之后做了对照实验:同一 IP 上"无 cookie → -352、只有 buvid3 → -352、buvid3+buvid4 → code:0",说明这是对本适配器主要失败模式的直接修复,而不是冗余保险 |
| 只用不带参数的 `v1/detail` | 加 `features=itemOpusStyle` | 用户实测反馈"有内容的动态没有 title":不带参数时 opus 图文帖返回 legacy 形态,`desc``null`,正文与标题整个丢失。带参数后同一 ID 返回 `major.opus.summary.text` / `title` / `pics`。AV / 转发帖不受影响,legacy 形态仍保留为兜底 |
---
## 4. 测试与验证
- 单元(13):正则匹配/拒绝/忽略短链、缓存键归一、图片映射(https 归一 + 缩略图 + `.gif → Animated`)、
封面、转发取 `orig` 媒体与正文拼接、纯文字无媒体、caption 转义、业务 code 分类(可重试性)、URL 归一、
设备 cookie 拼装。
- live3`#[ignore = "live network: …"]`):设备 cookie 可取、图片动态 2 图、纯文字动态无媒体。
CI 的 `live` job 已覆盖。动态接口被风控时这两条 live 测试打印 `skipping:` 并提前返回(与 pixiv 的
token 门控同款约定),设备 cookie 那条仍会真实执行。
- 实测命令:
`cargo run -p x-media --example fetch -- https://www.bilibili.com/opus/1245284537985925159`
(输出 2 张 `https://i0.hdslb.com/…jpg` + `@518w.jpg` 缩略图 + 话题 tags)。
- 全套:`cargo fmt --check``cargo clippy --workspace --all-targets -- -D warnings``cargo test --workspace` 全绿。
## 5. 已知限制
- 风控按 IP/请求量漂移,阶梯见 §2 最后一行:轻度靠设备 cookie 自愈,重度需 `BILIBILI_COOKIE` 或换 IP。
被拦时按**可重试**失败处理(队列退避)+ 一条 warn,不会静默丢帖。
- 接口 schema 会漂移(`module_dynamic.major` 实测可为 `null` 而正文留在 `desc`);DTO 全 `Option`
未知形态降级为"无媒体",不 panic。
- 动态内嵌视频只发封面图(与 bff 同策略),不下载流。
- 纯文字动态复用既有 "No media found" 回复。
- `b23.tv` 短链不被匹配(见上表)。
Generated
+154 -221
View File
@@ -10,25 +10,13 @@ checksum = "320119579fcad9c21884f5c4861d16174d0e06250625266f50fe6898340abefa"
[[package]]
name = "aes"
version = "0.8.4"
version = "0.9.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b169f7a6d4742236a0a00c541b845991d0ac43e546831af1249753ab4c3aa3a0"
checksum = "35f0f96ce78e38c3dc6d8948aa8163d06385be74000f3c7a95bf1eef35d3ea32"
dependencies = [
"cfg-if",
"cipher",
"cpufeatures 0.2.17",
]
[[package]]
name = "ahash"
version = "0.8.12"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5a15f179cd60c4584b8a8c596927aadc462e27f2ca70c04e0071964a73ba7a75"
dependencies = [
"cfg-if",
"once_cell",
"version_check",
"zerocopy",
"cpubits",
"cpufeatures",
]
[[package]]
@@ -72,15 +60,6 @@ dependencies = [
"object",
]
[[package]]
name = "arbitrary"
version = "1.4.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c3d036a3c4ab069c7b410a2ce876bd74808d2d0888a82667669f8e783a898bf1"
dependencies = [
"derive_arbitrary",
]
[[package]]
name = "atomic-waker"
version = "1.1.2"
@@ -171,11 +150,12 @@ checksum = "3ded4057c258ba199e2d26386d3af3780957ecaee6c4ef4041c6b4b8b97c0b06"
[[package]]
name = "block-buffer"
version = "0.10.4"
version = "0.12.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3078c7629b62d3f0439517fa394996acacc5cbc91c5a20d8c658e77abd503a71"
checksum = "d2f6c7dbe95a6ed67ad9f18e57daf93a2f034c524b99fd2b76d18fdfeb6660aa"
dependencies = [
"generic-array",
"hybrid-array",
"zeroize",
]
[[package]]
@@ -199,12 +179,6 @@ version = "1.25.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "95832e849adfb21180ccb6826a99da14e5d266ae5c2e668e1602cf234f153797"
[[package]]
name = "byteorder"
version = "1.5.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1fd0f2584146f6f2ef48085050886acf353beff7305ebd1ae69500e27c67f64b"
[[package]]
name = "bytes"
version = "1.12.1"
@@ -213,21 +187,11 @@ checksum = "fc652a48c352aef3ea3aed32080501cf3ef6ed5da78602a020c991775b0aff04"
[[package]]
name = "bzip2"
version = "0.5.2"
version = "0.6.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "49ecfb22d906f800d4fe833b6282cf4dc1c298f5057ca0b5445e5c209735ca47"
checksum = "f3a53fac24f34a81bc9954b5d6cfce0c21e18ec6959f44f56e8e90e4bb7c346c"
dependencies = [
"bzip2-sys",
]
[[package]]
name = "bzip2-sys"
version = "0.1.13+1.0.8"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "225bff33b2141874fe80d71e07d6eec4f85c5c216453dd96388240f96e1acc14"
dependencies = [
"cc",
"pkg-config",
"libbz2-rs-sys",
]
[[package]]
@@ -261,7 +225,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "65c35e4b699c7e15ccbe7ee35c005e4fc0a278d22238a2857e6ce2dadeda1b06"
dependencies = [
"cfg-if",
"cpufeatures 0.3.1",
"cpufeatures",
"rand_core 0.10.1",
]
@@ -279,14 +243,20 @@ dependencies = [
[[package]]
name = "cipher"
version = "0.4.4"
version = "0.5.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "773f3b9af64447d2ce9850330c473515014aa235e6a783b02db81ff39e4a3dad"
checksum = "e8cf2a2c93cd704877c0858356ed03480ff301ee950b43f1cbe4573b088bfa6c"
dependencies = [
"crypto-common",
"inout",
]
[[package]]
name = "cmov"
version = "0.5.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0c9ea0ac24bc397ab3c98583a3c9ba74fa56b09a4449bbe172b9b1ddb016027a"
[[package]]
name = "colored"
version = "3.1.1"
@@ -297,10 +267,16 @@ dependencies = [
]
[[package]]
name = "constant_time_eq"
version = "0.3.1"
name = "const-oid"
version = "0.10.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7c74b8349d32d297c9134b8c88677813a227df8f779daa29bfc29c183fe3dca6"
checksum = "a6ef517f0926dd24a1582492c791b6a4818a4d94e789a334894aa15b0d12f55c"
[[package]]
name = "constant_time_eq"
version = "0.4.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3d52eff69cd5e647efe296129160853a42795992097e8af39800e1060caeea9b"
[[package]]
name = "core-foundation-sys"
@@ -309,13 +285,10 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "773648b94d0e5d620f64f280777445740e61fe701025087ec8b57f45c791888b"
[[package]]
name = "cpufeatures"
version = "0.2.17"
name = "cpubits"
version = "0.1.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "59ed5838eebb26a2bb2e58f6d5b5316989ae9d08bab10e0e6d103e656d1b0280"
dependencies = [
"libc",
]
checksum = "15b85f9c39137c3a891689859392b1bd49812121d0d61c9caf00d46ed5ce06ae"
[[package]]
name = "cpufeatures"
@@ -326,21 +299,6 @@ dependencies = [
"libc",
]
[[package]]
name = "crc"
version = "3.4.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5eb8a2a1cd12ab0d987a5d5e825195d372001a4094a0376319d5a0ad71c1ba0d"
dependencies = [
"crc-catalog",
]
[[package]]
name = "crc-catalog"
version = "2.5.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "217698eaf96b4a3f0bc4f3662aaa55bdf913cd54d7204591faa790070c6d0853"
[[package]]
name = "crc32fast"
version = "1.5.2"
@@ -351,19 +309,21 @@ dependencies = [
]
[[package]]
name = "crossbeam-utils"
version = "0.8.23"
name = "crypto-common"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a31eee39dddec8330830986fcd7625edb5a24ec90ea038215273bbc3adb08ac6"
checksum = "ce6e4c961d6cd6c9a86db418387425e8bdeaf05b3c8bc1411e6dca4c252f1453"
dependencies = [
"hybrid-array",
]
[[package]]
name = "crypto-common"
version = "0.1.7"
name = "ctutils"
version = "0.4.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "78c8292055d1c1df0cce5d180393dc8cce0abec0a7102adb6c7b1eef6016d60a"
checksum = "7d5515a3834141de9eafb9717ad39eea8247b5674e6066c404e8c4b365d2a29e"
dependencies = [
"generic-array",
"typenum",
"cmov",
]
[[package]]
@@ -446,17 +406,6 @@ dependencies = [
"serde_core",
]
[[package]]
name = "derive_arbitrary"
version = "1.4.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1e567bd82dcff979e4b03460c307b3cdc9e96fde3d73bed1496d2bc75d9dd62a"
dependencies = [
"proc-macro2",
"quote",
"syn 2.0.119",
]
[[package]]
name = "derive_more"
version = "1.0.0"
@@ -480,13 +429,15 @@ dependencies = [
[[package]]
name = "digest"
version = "0.10.7"
version = "0.11.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9ed9a281f7bc9b7576e61468ba615a66a5c8cfdff42420a70aa82701a3b1e292"
checksum = "f1dd6dbb5841937940781866fa1281a1ff7bd3bf827091440879f9994983d5c2"
dependencies = [
"block-buffer",
"const-oid",
"crypto-common",
"subtle",
"ctutils",
"zeroize",
]
[[package]]
@@ -632,6 +583,12 @@ dependencies = [
"zlib-rs",
]
[[package]]
name = "foldhash"
version = "0.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "77ce24cb58228fbb8aa041425bb1050850ac19177686ea6e0f41a70416f56fdb"
[[package]]
name = "form_urlencoded"
version = "1.2.2"
@@ -729,16 +686,6 @@ dependencies = [
"slab",
]
[[package]]
name = "generic-array"
version = "0.14.7"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "85649ca51fd72272d7821adaf274ad91c288277713d9c18820d8499a7ff69e9a"
dependencies = [
"typenum",
"version_check",
]
[[package]]
name = "getrandom"
version = "0.2.17"
@@ -752,20 +699,6 @@ dependencies = [
"wasm-bindgen",
]
[[package]]
name = "getrandom"
version = "0.3.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "899def5c37c4fd7b2664648c28120ecec138e4d395b459e5ca34f9cce2dd77fd"
dependencies = [
"cfg-if",
"js-sys",
"libc",
"r-efi 5.3.0",
"wasip2",
"wasm-bindgen",
]
[[package]]
name = "getrandom"
version = "0.4.3"
@@ -775,7 +708,7 @@ dependencies = [
"cfg-if",
"js-sys",
"libc",
"r-efi 6.0.0",
"r-efi",
"rand_core 0.10.1",
"wasm-bindgen",
]
@@ -788,11 +721,11 @@ checksum = "8a9ee70c43aaf417c914396645a0fa852624801b24ebb7ae78fe8272889ac888"
[[package]]
name = "hashbrown"
version = "0.14.5"
version = "0.16.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e5274423e17b7c9fc20b6e7e208532f9b19825d82dfd615708b70edd83df41f1"
checksum = "841d1cc9bed7f9236f321df977030373f4a4163ae1a7dbfe1a51a2c1a51d9100"
dependencies = [
"ahash",
"foldhash",
]
[[package]]
@@ -800,14 +733,17 @@ name = "hashbrown"
version = "0.17.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ed5909b6e89a2db4456e54cd5f673791d7eca6732202bbf2a9cc504fe2f9b84a"
dependencies = [
"foldhash",
]
[[package]]
name = "hashlink"
version = "0.9.1"
version = "0.12.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6ba4ff7128dee98c7dc9794b6a411377e1404dba1c97deb8d1a55297bd25d8af"
checksum = "a596f1b20ed2cc5ecac41a164aaebc7258057060f06c0cf7a2ba3991ee7990fb"
dependencies = [
"hashbrown 0.14.5",
"hashbrown 0.17.1",
]
[[package]]
@@ -830,9 +766,9 @@ checksum = "7f24254aa9a54b5c858eaee2f5bccdb46aaf0e486a595ed5fd8f86ba55232a70"
[[package]]
name = "hmac"
version = "0.12.1"
version = "0.13.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6c49c37c09c17a53d937dfbb742eb3a961d65a994e6bcdcf37e7399d0cc8ab5e"
checksum = "6303bc9732ae41b04cb554b844a762b4115a61bfaa81e3e83050991eeb56863f"
dependencies = [
"digest",
]
@@ -894,6 +830,15 @@ version = "2.4.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "15cdd26707701c53297e2fa6afb323d55fbc1d0810c3aec078ae3ef0424c3c15"
[[package]]
name = "hybrid-array"
version = "0.4.15"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "27f864f10dfb56725ce5ce5472bc52252c8f93a4ab86327122cebf62c5f59a17"
dependencies = [
"typenum",
]
[[package]]
name = "hyper"
version = "1.11.1"
@@ -1132,11 +1077,11 @@ dependencies = [
[[package]]
name = "inout"
version = "0.1.4"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "879f10e63c20629ecabbb64a8010319738c66a5cd0c29b02d63d272b03751d01"
checksum = "4250ce6452e92010fdf7268ccc5d14faa80bb12fc741938534c58f16804e03c7"
dependencies = [
"generic-array",
"hybrid-array",
]
[[package]]
@@ -1252,6 +1197,12 @@ dependencies = [
"wasm-bindgen",
]
[[package]]
name = "libbz2-rs-sys"
version = "0.2.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "34b357333733e8260735ba5894eb928c02ecc69c78715f01a8019e7fa7f2db4c"
[[package]]
name = "libc"
version = "0.2.189"
@@ -1260,9 +1211,9 @@ checksum = "3eaf3ede3fee6db1a4c2ee091bf8a8b4dccdc6d17f656fb07896ee72867612f2"
[[package]]
name = "libsqlite3-sys"
version = "0.30.1"
version = "0.38.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2e99fb7a497b1e3339bc746195567ed8d3e24945ecd636e3619d20b9de9e9149"
checksum = "f1d20bef17f513b9b3004532233187769cd072d790971f4e4da0e346eb6401e8"
dependencies = [
"cc",
"pkg-config",
@@ -1309,24 +1260,12 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "4050469837a6ff301cd14c1f8f24f88549e6d548f24f64e2148eb0f72cebc51f"
[[package]]
name = "lzma-rs"
version = "0.3.0"
name = "lzma-rust2"
version = "0.16.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "297e814c836ae64db86b36cf2a557ba54368d03f6afcd7d947c266692f71115e"
checksum = "ca93e534d1142d1d0dcca6d25fe302508a5dfb40b302802904577725ea0b695b"
dependencies = [
"byteorder",
"crc",
]
[[package]]
name = "lzma-sys"
version = "0.1.20"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5fda04ab3764e6cde78b9974eec4f779acaba7c4e84b36eca3cf77c581b85d27"
dependencies = [
"cc",
"libc",
"pkg-config",
"sha2",
]
[[package]]
@@ -1443,9 +1382,9 @@ dependencies = [
[[package]]
name = "pbkdf2"
version = "0.12.2"
version = "0.13.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f8ed6a7761f76e3b9f92dfb0a60a6a6477c61024b775147ff0973a02653abaf2"
checksum = "112d82ceb8c5bf524d9af484d4e4970c9fd5a0cc15ba14ad93dccd28873b0629"
dependencies = [
"digest",
"hmac",
@@ -1532,6 +1471,12 @@ version = "0.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "439ee305def115ba05938db6eb1644ff94165c5ab5e9420d1c1bcedbba909391"
[[package]]
name = "ppmd-rust"
version = "1.5.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "196a7c80b9a7652aba7cc070827516c2abe4ccdf53d128e1944003cf5726cff1"
[[package]]
name = "ppv-lite86"
version = "0.2.21"
@@ -1656,12 +1601,6 @@ dependencies = [
"proc-macro2",
]
[[package]]
name = "r-efi"
version = "5.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "69cdb34c158ceb288df11e18b4bd39de994f6657d83847bdffdbd7f346754b0f"
[[package]]
name = "r-efi"
version = "6.0.0"
@@ -1857,10 +1796,20 @@ dependencies = [
]
[[package]]
name = "rusqlite"
version = "0.32.1"
name = "rsqlite-vfs"
version = "0.1.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7753b721174eb8ff87a9a0e799e2d7bc3749323e773db92e0984debb00019d6e"
checksum = "c51c9ae4df8a7fba42103df5c621fa3c37eccf3a3c650879e90fc48b11cc192c"
dependencies = [
"hashbrown 0.16.1",
"thiserror",
]
[[package]]
name = "rusqlite"
version = "0.40.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "23f2a97da3e3873c73cb2a2e71b35c40ff95e0b1eefa8d72d8499a6928c3b5b3"
dependencies = [
"bitflags 2.13.2",
"fallible-iterator",
@@ -1868,6 +1817,7 @@ dependencies = [
"hashlink",
"libsqlite3-sys",
"smallvec",
"sqlite-wasm-rs",
]
[[package]]
@@ -2067,12 +2017,23 @@ dependencies = [
[[package]]
name = "sha1"
version = "0.10.7"
version = "0.11.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a978451301f4db1d02937a4ab3ccce137717b81826e79b7d49ffe3244a13c3b8"
checksum = "aacc4cc499359472b4abe1bf11d0b12e688af9a805fa5e3016f9a386dc2d0214"
dependencies = [
"cfg-if",
"cpufeatures 0.2.17",
"cpufeatures",
"digest",
]
[[package]]
name = "sha2"
version = "0.11.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "446ba717509524cb3f22f17ecc096f10f4822d76ab5c0b9822c5f9c284e825f4"
dependencies = [
"cfg-if",
"cpufeatures",
"digest",
]
@@ -2120,6 +2081,18 @@ dependencies = [
"windows-sys 0.61.2",
]
[[package]]
name = "sqlite-wasm-rs"
version = "0.5.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dc3efc0da82635d7e1ced0053bbbfa8c7ab9645d0bf36ceb4f7127bb85315d75"
dependencies = [
"cc",
"js-sys",
"rsqlite-vfs",
"wasm-bindgen",
]
[[package]]
name = "stable_deref_trait"
version = "1.2.1"
@@ -2328,6 +2301,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cdb87b95ec50ddfa440816d227a17b2ccbdda963a316a727fda0fc4334f7d134"
dependencies = [
"deranged",
"js-sys",
"num-conv",
"powerfmt",
"serde_core",
@@ -2502,6 +2476,12 @@ version = "0.2.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e421abadd41a4225275504ea4d6566923418b7f05506fbc9c0fe86ba7396114b"
[[package]]
name = "typed-path"
version = "0.12.3"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8e28f89b80c87b8fb0cf04ab448d5dd0dd0ade2f8891bae878de66a75a28600e"
[[package]]
name = "typenum"
version = "1.20.1"
@@ -2568,12 +2548,6 @@ version = "0.2.15"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "accd4ea62f7bb7a82fe23066fb0957d48ef677f6eeb8215f372f52e48bb32426"
[[package]]
name = "version_check"
version = "0.9.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0b928f33d975fc6ad9f86c8f283853ad26bdd5b10b7f1542aa2fa15e2289105a"
[[package]]
name = "want"
version = "0.3.1"
@@ -2589,15 +2563,6 @@ version = "0.11.1+wasi-snapshot-preview1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ccf3ec651a847eb01de73ccad15eb7d99f80485de043efb2f370cd654f4ea44b"
[[package]]
name = "wasip2"
version = "1.0.4+wasi-0.2.12"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b67efb37e106e55ce722a510d6b5f9c17f083e5fc79afc2badeb12cc313d9487"
dependencies = [
"wit-bindgen",
]
[[package]]
name = "wasm-bindgen"
version = "0.2.128"
@@ -2845,12 +2810,6 @@ version = "0.52.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "589f6da84c646204747d1270a2a5661ea66ed1cced2631d546fdfb155959f9ec"
[[package]]
name = "wit-bindgen"
version = "0.57.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1ebf944e87a7c253233ad6766e082e3cd714b5d03812acc24c318f549614536e"
[[package]]
name = "writeable"
version = "0.6.4"
@@ -2859,13 +2818,13 @@ checksum = "3ad82d2a33cdc9674dc7465672f271e096168fcdbe0f799d9e6db8c5892679dc"
[[package]]
name = "x-media"
version = "1.6.0"
version = "1.8.0"
dependencies = [
"bytes",
"dotenv",
"html-escape",
"log",
"rand 0.8.8",
"rand 0.10.2",
"regex",
"reqwest",
"serde",
@@ -2879,7 +2838,7 @@ dependencies = [
[[package]]
name = "xmedia-bot"
version = "1.6.0"
version = "1.8.0"
dependencies = [
"bytes",
"dotenv",
@@ -2890,7 +2849,7 @@ dependencies = [
"parking_lot",
"png",
"pretty_env_logger",
"rand 0.8.8",
"rand 0.10.2",
"rusqlite",
"serde",
"serde_json",
@@ -2902,15 +2861,6 @@ dependencies = [
"zune-jpeg",
]
[[package]]
name = "xz2"
version = "0.1.7"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "388c44dc09d76f1536602ead6d325eb532f5c122f17782bd57fb47baeeb767e2"
dependencies = [
"lzma-sys",
]
[[package]]
name = "yoke"
version = "0.8.3"
@@ -2980,20 +2930,6 @@ name = "zeroize"
version = "1.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e13c156562582aa81c60cb29407084cdb54c4164760106ab78e6c5b0858cf64e"
dependencies = [
"zeroize_derive",
]
[[package]]
name = "zeroize_derive"
version = "1.5.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3c50655cbb0fe3fc43170059e702f1ce5e19b84cec58dc87b037a09935c2f328"
dependencies = [
"proc-macro2",
"quote",
"syn 2.0.119",
]
[[package]]
name = "zerotrie"
@@ -3030,29 +2966,26 @@ dependencies = [
[[package]]
name = "zip"
version = "2.4.2"
version = "8.6.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "fabe6324e908f85a1c52063ce7aa26b68dcb7eb6dbc83a2d148403c9bc3eba50"
checksum = "2d04a6b5381502aa6087c94c669499eb1602eb9c5e8198e534de571f7154809b"
dependencies = [
"aes",
"arbitrary",
"bzip2",
"constant_time_eq",
"crc32fast",
"crossbeam-utils",
"deflate64",
"displaydoc",
"flate2",
"getrandom 0.3.4",
"getrandom 0.4.3",
"hmac",
"indexmap 2.14.2",
"lzma-rs",
"lzma-rust2",
"memchr",
"pbkdf2",
"ppmd-rust",
"sha1",
"thiserror",
"time",
"xz2",
"typed-path",
"zeroize",
"zopfli",
"zstd",
+23 -13
View File
@@ -1,14 +1,18 @@
# TelegramXMediaBot
A Telegram bot that turns post links from X / Twitter, Pixiv, Bluesky, and Misskey (misskey.io) into media messages (images, video, GIF) with the post's title, author, and tags.
A Telegram bot that turns post links from X / Twitter, Pixiv, Bluesky, Misskey (misskey.io), and Bilibili dynamics into media messages (images, video, GIF) with the post's title, author, and tags.
## Features
- Sending a link in a private chat fetches and sends the images, videos and GIFs automatically; oversized media is split into batches
- Text-only posts report "no media"; unsupported links are silently ignored
- Inline queries (`@bot <link>`)
- Bind a forward channel for automatic forwarding; edit the caption before forwarding and apply custom templates
- Failed sends are retried automatically with persistence; the user is notified after retries are exhausted
- Sending a link in a private chat fetches and sends the images, videos and GIFs automatically; oversized media is split into batches (10 items per group)
- Text-only posts report "no media"; unsupported links are silently ignored. Fetch failures name the reason (post gone / content withheld / source risk control / site not enabled)
- Long posts (text ≥ `CAPTION_QUOTE_TEXT_CHARS`, default 200) show **the text part** of their caption inside a collapsible blockquote, with the link and author line left outside it
- Inline queries (`@bot <link>`) — except Pixiv images and locally transcoded animations, which Telegram cannot fetch (no Referer) and would show broken, so they are skipped; a supported link posted in a group gets a one-line hint to use the private chat or inline mode (channels stay silent)
- `/start` explains the supported sites and how to use it; `/help` lists the commands plus argument syntax, the caption placeholders and the private-chat rule; the bot's profile description texts are set at startup
- `/settings` shows this chat's configuration (forward channel, edit-before-forward, per-site caption formats, saved templates); templates are added with `/set_template` and removed with `/remove_template`
- Bind a forward channel for automatic forwarding; edit the caption before forwarding and apply custom templates (the prompt carries Confirm / Skip buttons, states its expiry, and is marked expired in place once it lapses)
- Failed sends are retried automatically with persistence; the notice names which link failed, how long the retry waits, or the final cause
- The chat action stays on screen for the whole fetch, so long jobs (ugoira transcode, large uploads) do not look stalled
- Pixiv ugoira animations are transcoded to MP4; Bluesky videos are remuxed (HLS stream → MP4)
- Photos exceeding Telegram's size/dimension limits are compressed automatically (original format kept, JPEG fallback only when needed)
- Link-result cache: after a successful send the Telegram file ids and caption fields are cached locally, so a repeated link is re-sent from local state — no source-site request, no media file stored (expiry controlled by `LINK_CACHE_TTL_SECONDS`, default 7 days)
@@ -30,9 +34,11 @@ docker build -t tgxmb .
docker run --rm -d --name tgxmb --env-file .env -v ./data:/app/data tgxmb
```
Environment variables: `TELOXIDE_TOKEN` (required), `PIXIV_REFRESH_TOKEN`, `BOT_ADMIN`, `EDIT_MESSAGE_TTL_SECONDS`, `LINK_CACHE_TTL_SECONDS`, `RUST_LOG`, `TELOXIDE_PROXY`, `WEBHOOK*`, `TWITTER_AUTH_TOKEN` (optional).
Environment variables: `TELOXIDE_TOKEN` (required), `PIXIV_REFRESH_TOKEN`, `BOT_ADMIN`, `EDIT_MESSAGE_TTL_SECONDS`, `LINK_CACHE_TTL_SECONDS`, `RUST_LOG`, `TELOXIDE_PROXY`, `WEBHOOK*`, `TWITTER_AUTH_TOKEN` (optional), `BILIBILI_COOKIE` (optional).
NSFW tweets: the public syndication endpoint does not return sensitive content. Setting `TWITTER_AUTH_TOKEN` (the `auth_token` cookie value of a logged-in x.com session) lets the bot fetch NSFW media in the logged-in state only when it hits a withheld tweet; without it, the bot reports no media.
NSFW tweets: the public syndication endpoint does not return sensitive content. Setting `TWITTER_AUTH_TOKEN` (the `auth_token` cookie value of a logged-in x.com session) lets the bot fetch NSFW media in the logged-in state only when it hits a withheld tweet; without it the bot answers that the post's media is withheld and needs `TWITTER_AUTH_TOKEN`.
Bilibili dynamics are fetched anonymously by default (no login; the bot fetches bilibili's anonymous `buvid3`/`buvid4` device cookies itself to raise the success rate). If the server's egress IP gets hard-flagged by bilibili (persistent `risk control (-352)` log lines or HTTP 412), set `BILIBILI_COOKIE` (the whole cookie string from a logged-in browser, e.g. `SESSDATA=…; bili_jct=…`) to restore access. Only a dynamic's images and animations are sent; an attached video degrades to its cover image.
### Webhook deployment (needs a reverse proxy)
@@ -80,10 +86,12 @@ Telegram only accepts ports 443/80/88/8443.
| Variable | Description |
|---|---|
| `TELOXIDE_TOKEN` | Bot token (required) |
| `PIXIV_REFRESH_TOKEN` | Pixiv refresh token; Pixiv is disabled without it |
| `PIXIV_REFRESH_TOKEN` | Pixiv refresh token; Pixiv is disabled without it (a pixiv link then gets an explicit "site not enabled" reply instead of silence) |
| `BILIBILI_COOKIE` | Optional bilibili cookie string (`SESSDATA=…; bili_jct=…`); only needed when the egress IP stays risk-controlled (device cookies are fetched automatically) |
| `BOT_ADMIN` | Admin chat IDs, comma-separated; receives start/stop notifications |
| `EDIT_MESSAGE_TTL_SECONDS` | Edit-before-forward record expiry in seconds, default 86400 |
| `EDIT_MESSAGE_TTL_SECONDS` | Edit-before-forward record expiry in seconds, default 86400; once lapsed the prompt is rewritten in place to "expired — nothing was forwarded" (no extra message) |
| `LINK_CACHE_TTL_SECONDS` | Link-result cache expiry in seconds, default 604800 (7 days) |
| `CAPTION_QUOTE_TEXT_CHARS` | **The text part** of the caption (the joined `{title}` + `{content}`) is wrapped in a collapsible blockquote once it reaches this many characters, default 200; `0` disables |
| `DATA_DIR` | Data directory (where the SQLite `task_queue.db` lives), default `data` (relative to the working directory, created automatically) |
| `RUST_LOG` | Log level |
| `TELOXIDE_PROXY` | HTTP proxy (e.g. `http://127.0.0.1:10808`); applies to both the Telegram Bot API and site fetches — required on restricted networks (e.g. behind the GFW) |
@@ -109,15 +117,17 @@ Telegram only accepts ports 443/80/88/8443.
| `/help` | List all commands and usage (this command table) |
| `/set_forward_channel <channel>` | Set the forward channel: `@channel` or channel ID; media messages are forwarded to it automatically afterwards |
| `/remove_forward_channel` | Remove the forward channel |
| `/edit_before_forward` | Toggle "edit before forward": when enabled, the bot posts a prompt after forwarding; replying to it edits the first forwarded message's caption (or taps a template button to apply one) |
| `/edit_before_forward` | Toggle "edit before forward": when enabled, the bot posts a prompt after forwarding; replying to it edits the first forwarded message's caption (or tapping a template button applies one), then `↩️ Confirm` forwards and `🛑 Skip` drops this forward; the prompt states its expiry and is marked expired in place when it lapses (nothing is forwarded) |
| `/set_template <name>` | Reply to a message containing `[]` to save it as a named template; `[]` is replaced by the original post link when forwarding (used with "edit before forward") |
| `/set_format <site> <format>` | Customize the caption format for one site. Sites: `twitter` / `bsky` / `pixiv` / `misskey`. Placeholders: `{url}` `{author}` `{author_url}` `{title}` `{tags}` |
| `/remove_template <name>` | Remove a template (names are listed by `/settings`; the prompt's keyboard shows at most 60) |
| `/settings` | Show this chat's configuration: forward channel, edit-before-forward, per-site caption formats, saved templates |
| `/set_format <site> <format>` | Customize the caption format for one site. Sites: `twitter` / `bsky` / `pixiv` / `misskey` / `bilibili`. Placeholders: `{url}` `{author}` `{author_url}` `{title}` `{content}` `{tags}`; unknown placeholders are rejected with the list of valid ones, and `-` restores the site's built-in format (preview with `/debug <link>`) |
| `/clear_cache [link]` | Clear the link cache (admin only); with a link only that entry, otherwise everything |
| `/bot_dict` | Show the current chat state (debugging; admin only) |
| `/test <link>` | Parse a link and send its media; no channel forward, no edit-before-forward prompt (send only) |
| `/debug <link>` | Debug: parse a link and report the parse result only (site, title, author, tags, media list) — no media is sent |
Link processing works only in private chats; commands work in any chat.
Link processing works only in private chats; commands work in any chat. A supported link posted in a group gets a one-line hint to use the private chat or inline mode; channels stay silent.
## Notes
+23 -13
View File
@@ -1,14 +1,18 @@
# TelegramXMediaBot
Telegram 机器人,将 X / Twitter、Pixiv、Bluesky、Misskey (misskey.io) 的帖子链接转换为媒体消息发送,附带帖子标题、作者与标签。
Telegram 机器人,将 X / Twitter、Pixiv、Bluesky、Misskey (misskey.io)、Bilibili 动态的帖子链接转换为媒体消息发送,附带帖子标题、作者与标签。
## 功能
- 私聊发送链接后自动抓取并发送图片、视频与 GIF,超量图片自动分批
- 纯文字帖提示无媒体;不支持的链接静默忽略
- 支持内联查询(`@机器人 <链接>`
- 可绑定转发频道自动转发;支持转发前编辑 caption 与自定义模板
- 发送失败自动重试并持久化,重试耗尽后通知用户
- 私聊发送链接后自动抓取并发送图片、视频与 GIF,超量图片自动分批(每批 10 张)
- 纯文字帖提示无媒体;不支持的链接静默忽略。抓取失败会按原因分别提示(帖子已删除 / 内容受限 / 源站风控 / 站点未启用)
- 长帖(正文 ≥ `CAPTION_QUOTE_TEXT_CHARS`,默认 200)的**正文部分**用可折叠引用块展示,链接与作者行留在引用块外
- 支持内联查询(`@机器人 <链接>`;Pixiv 图片与本地转码的动图不支持内联 —— Telegram 取图时无法携带 Referer,会显示破图,因此跳过);在群聊里发链接会提示改用私聊或内联查询(频道内保持静默)
- `/start` 说明支持的站点与用法,`/help` 列出命令、参数格式、caption 占位符与私聊限制;bot 资料页(description / short description)启动时一并设置
- `/settings` 查看本聊天配置(转发频道、转发前编辑开关、各站点 caption 格式、模板列表);模板可用 `/set_template` 增、`/remove_template`
- 可绑定转发频道自动转发;支持转发前编辑 caption 与自定义模板(提示消息带 Confirm / Skip 按钮并写明过期时间,过期后就地标记为已过期)
- 发送失败自动重试并持久化,重试耗尽后通知用户;提示会写明是哪条链接、重试等待多久、或最终失败的原因
- 抓取期间持续显示"正在输入 / 正在发送"状态,长任务(ugoira 转码、大图上传)不会看起来卡死
- Pixiv ugoira 动图自动转码为 MP4Bluesky 视频自动转码(HLS 流 → MP4)
- 超过 Telegram 尺寸/大小限制的图片自动压缩(保持原格式,必要时转 JPEG)
- 链接结果本地缓存:成功发送后缓存 Telegram file id 与 caption 等,再次收到相同链接直接本地重发,不再请求源站、不保存媒体文件(`LINK_CACHE_TTL_SECONDS` 控制过期,默认 7 天)
@@ -30,9 +34,11 @@ docker build -t tgxmb .
docker run --rm -d --name tgxmb --env-file .env -v ./data:/app/data tgxmb
```
环境变量:`TELOXIDE_TOKEN`(必填)、`PIXIV_REFRESH_TOKEN``BOT_ADMIN``EDIT_MESSAGE_TTL_SECONDS``LINK_CACHE_TTL_SECONDS``RUST_LOG``TELOXIDE_PROXY``WEBHOOK*``TWITTER_AUTH_TOKEN`(可选)。
环境变量:`TELOXIDE_TOKEN`(必填)、`PIXIV_REFRESH_TOKEN``BOT_ADMIN``EDIT_MESSAGE_TTL_SECONDS``LINK_CACHE_TTL_SECONDS``RUST_LOG``TELOXIDE_PROXY``WEBHOOK*``TWITTER_AUTH_TOKEN`(可选)`BILIBILI_COOKIE`(可选)
NSFW 推文:公开的 syndication 接口不返回敏感内容。设置 `TWITTER_AUTH_TOKEN`(登录 x.com 后浏览器 Cookie 里的 `auth_token` 值)后,bot 会仅在遇到 NSFW 推文时以登录态获取媒体;未设置则提示无媒体
NSFW 推文:公开的 syndication 接口不返回敏感内容。设置 `TWITTER_AUTH_TOKEN`(登录 x.com 后浏览器 Cookie 里的 `auth_token` 值)后,bot 会仅在遇到 NSFW 推文时以登录态获取媒体;未设置则回复该推文内容受限(需要配置 `TWITTER_AUTH_TOKEN`
Bilibili 动态默认匿名抓取(无需登录,bot 会自动从 B 站的匿名指纹接口取 `buvid3`/`buvid4` 设备 cookie 以提高成功率)。若服务器出口 IP 被 B 站重度风控(日志里的 `risk control (-352)` 或 HTTP 412,且持续出现),设置 `BILIBILI_COOKIE`(登录后浏览器里整条 Cookie 串,如 `SESSDATA=…; bili_jct=…`)可恢复访问。当前只发送动态里的图片与动图,动态内嵌视频发送其封面。
### Webhook 部署(需要反向代理)
@@ -80,10 +86,12 @@ Telegram 只接受 443/80/88/8443 端口。
| 变量 | 说明 |
|---|---|
| `TELOXIDE_TOKEN` | Bot token(必填) |
| `PIXIV_REFRESH_TOKEN` | Pixiv 刷新令牌;未设置则禁用 Pixiv |
| `PIXIV_REFRESH_TOKEN` | Pixiv 刷新令牌;未设置则禁用 Pixiv(此时收到 pixiv 链接会明确回复「站点未启用」,不会静默忽略) |
| `BILIBILI_COOKIE` | 可选的 B 站 Cookie 串(`SESSDATA=…; bili_jct=…`),仅在出口 IP 被持续风控时才需要(设备 cookie 由 bot 自动获取) |
| `BOT_ADMIN` | 管理员聊天 ID,逗号分隔;接收启动/停止通知 |
| `EDIT_MESSAGE_TTL_SECONDS` | 转发前编辑记录过期秒数,默认 86400 |
| `EDIT_MESSAGE_TTL_SECONDS` | 转发前编辑记录过期秒数,默认 86400;过期后提示消息会被就地改写为「已过期,未转发」(不额外发消息打扰) |
| `LINK_CACHE_TTL_SECONDS` | 链接结果缓存过期秒数,默认 604800(7 天) |
| `CAPTION_QUOTE_TEXT_CHARS` | 正文(`{title}` + `{content}` 合计)达到该长度(字符)时,caption 的**正文部分**用可折叠引用块包裹,默认 200;`0` 关闭 |
| `DATA_DIR` | 数据目录(SQLite 数据库 `task_queue.db` 所在目录),默认 `data`(相对工作目录,会自动创建) |
| `RUST_LOG` | 日志级别 |
| `TELOXIDE_PROXY` | HTTP 代理(如 `http://127.0.0.1:10808`);同时作用于 Telegram Bot API 与站点抓取请求,网络受限环境(如 GFW)必需 |
@@ -109,15 +117,17 @@ Telegram 只接受 443/80/88/8443 端口。
| `/help` | 查看全部命令及用法(即本文档的命令表) |
| `/set_forward_channel <频道>` | 设置转发频道,参数为 `@频道名` 或频道 ID;设置后发送的媒体消息会自动转发到该频道 |
| `/remove_forward_channel` | 取消转发频道 |
| `/edit_before_forward` | 开关「转发前编辑」:开启后,转发成功后 bot 会发一条提示消息,回复它可修改第一条转发消息的 caption(或点击模板按钮套用模板) |
| `/edit_before_forward` | 开关「转发前编辑」:开启后,转发成功后 bot 会发一条提示消息,回复它可修改第一条转发消息的 caption(或点击模板按钮套用模板),再点 `↩️ Confirm` 才会真正转发,`🛑 Skip` 放弃本次转发;提示消息写明过期时间,过期后原地标记为已过期且不会转发 |
| `/set_template <名称>` | 回复一条含 `[]` 的消息,将其保存为命名模板;转发时 `[]` 会被替换为原帖链接(配合「转发前编辑」使用) |
| `/set_format <站点> <格式>` | 自定义某站点的 caption 格式。站点:`twitter` / `bsky` / `pixiv` / `misskey`。占位符:`{url}` `{author}` `{author_url}` `{title}` `{tags}` |
| `/remove_template <名称>` | 删除某个模板(名称见 `/settings`;提示消息的模板按钮最多显示 60 个) |
| `/settings` | 查看本聊天配置:转发频道、转发前编辑开关、各站点 caption 格式、模板列表 |
| `/set_format <站点> <格式>` | 自定义某站点的 caption 格式。站点:`twitter` / `bsky` / `pixiv` / `misskey` / `bilibili`。占位符:`{url}` `{author}` `{author_url}` `{title}` `{content}` `{tags}`;未识别的占位符会被拒绝并列出可用项,格式填 `-` 恢复站点默认格式(可用 `/debug <链接>` 预览效果) |
| `/clear_cache [链接]` | 清空链接缓存(仅管理员);带链接只清该条,否则清空全部 |
| `/bot_dict` | 查看当前聊天状态(调试用;仅管理员) |
| `/test <链接>` | 解析链接并发送媒体;不转发到频道、不弹转发前编辑提示(仅发送) |
| `/debug <链接>` | 调试:只解析链接并返回解析结果(站点、标题、作者、标签、媒体列表),不发送任何媒体 |
链接处理仅限私聊;命令在任意聊天可用。
链接处理仅限私聊;命令在任意聊天可用。在群聊里发受支持的链接会回复一条提示(改用私聊或内联查询),频道内保持静默。
## 备注
+3 -3
View File
@@ -1,6 +1,6 @@
[package]
name = "x-media"
version = "1.6.0"
version = "1.8.0"
edition = "2024"
[dependencies]
@@ -11,10 +11,10 @@ regex = "1.12"
html-escape = "0.2"
url = "2.5.2"
bytes = "1"
zip = "2"
zip = "8"
tempfile = "3"
thiserror = "2"
rand = "0.8"
rand = "0.10"
log = "0.4"
tokio = { version = "1.40", features = ["time"] }
File diff suppressed because it is too large Load Diff
+6
View File
@@ -0,0 +1,6 @@
mod interface;
mod model;
pub use interface::{
BilibiliSite, PATTERN, cache_key, enabled, fetch_from_url, is_retryable, media_headers,
};
+158
View File
@@ -0,0 +1,158 @@
//! Serde DTOs for the Bilibili dynamic detail endpoint
//! (`/x/polymer/web-dynamic/v1/detail`), mirroring live responses
//! (field paths verified 2026-09-17). Every field is optional so an API
//! shape change degrades to "no media" instead of a parse failure.
use serde::Deserialize;
#[derive(Deserialize, Debug)]
pub(crate) struct Detail {
/// Business code: `0` = OK, `-352`/`-412` = risk control, `500`/`4101147`
/// = gone.
pub(crate) code: i64,
#[serde(default)]
pub(crate) message: Option<String>,
#[serde(default)]
pub(crate) data: Option<Data>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Data {
#[serde(default)]
pub(crate) item: Option<Box<Item>>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Item {
/// The dynamic id, same numeric id as in the URL.
#[serde(default)]
pub(crate) id_str: String,
#[serde(default)]
pub(crate) modules: Option<Modules>,
/// The quoted dynamic when this item is a forward. A forward shell often
/// carries no media of its own — the original holds it.
#[serde(default)]
pub(crate) orig: Option<Box<Item>>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Modules {
#[serde(default)]
pub(crate) module_author: Option<Author>,
#[serde(default)]
pub(crate) module_dynamic: Option<Dynamic>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Author {
#[serde(default)]
pub(crate) name: String,
#[serde(default)]
pub(crate) mid: Option<i64>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Dynamic {
#[serde(default)]
pub(crate) desc: Option<Desc>,
#[serde(default)]
pub(crate) major: Option<Major>,
/// A single topic (`{"id":…,"name":…}`), the dynamic's only tag source.
#[serde(default)]
pub(crate) topic: Option<Topic>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Desc {
#[serde(default)]
pub(crate) text: String,
}
/// `major` is a tagged union: `type` (`MAJOR_TYPE_DRAW` / `_OPUS` /
/// `_ARCHIVE` / …) plus one payload object per type. Only the three payloads
/// this adapter reads are modeled; an unknown major simply yields no media.
#[derive(Deserialize, Debug)]
pub(crate) struct Major {
#[serde(default)]
pub(crate) draw: Option<Draw>,
#[serde(default)]
pub(crate) opus: Option<Opus>,
#[serde(default)]
pub(crate) archive: Option<Archive>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Draw {
#[serde(default)]
pub(crate) items: Vec<Pic>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Pic {
/// `major.draw` image URL.
#[serde(default)]
pub(crate) src: Option<String>,
/// `major.opus.pics` image URL — the opus shape names the field
/// differently while carrying the same image.
#[serde(default)]
pub(crate) url: Option<String>,
}
impl Pic {
/// The image URL, whichever key this serialization put it under.
pub(crate) fn url(&self) -> Option<&str> {
self.src.as_deref().or(self.url.as_deref())
}
}
/// `major.opus`: the serialization of an image/text post the web client asks
/// for (`features=itemOpusStyle`). It carries the parts the legacy shape drops
/// entirely — the document title and body of an opus post, whose
/// `module_dynamic.desc` comes back `null`.
#[derive(Deserialize, Debug)]
pub(crate) struct Opus {
/// Document headline; often absent.
#[serde(default)]
pub(crate) title: Option<String>,
/// Document body (untruncated: a 307-char sample came back whole).
#[serde(default)]
pub(crate) summary: Option<Desc>,
#[serde(default)]
pub(crate) pics: Vec<Pic>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Archive {
/// The attached video's cover — the only image an AV dynamic has (the
/// video itself is deliberately not resolved, see the module docs).
#[serde(default)]
pub(crate) cover: Option<String>,
/// The video's title. An AV dynamic has no body of its own (`desc` comes
/// back `null`), so this card title is the post's content.
#[serde(default)]
pub(crate) title: Option<String>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Topic {
#[serde(default)]
pub(crate) name: String,
}
/// Response of the anonymous fingerprint endpoint (`/x/frontend/finger/spi`),
/// the source of the adapter's device cookies.
#[derive(Deserialize, Debug)]
pub(crate) struct Fingerprint {
#[serde(default)]
pub(crate) data: Option<FingerprintData>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct FingerprintData {
/// Sent as the `buvid3` cookie.
#[serde(default, rename = "b_3")]
pub(crate) buvid3: String,
/// Sent as the `buvid4` cookie.
#[serde(default, rename = "b_4")]
pub(crate) buvid4: String,
}
+7 -3
View File
@@ -330,13 +330,16 @@ impl From<Post> for Fetched {
url: url.clone(),
author: encode_text(&post.author).into_owned(),
author_url: author_url.clone(),
title: encode_text(&post.text).into_owned(),
// A post has no title: its text is all content.
title: String::new(),
content: encode_text(&post.text).into_owned(),
tags: String::new(),
});
Fetched {
source_url: url,
caption: post.caption(),
title: post.text.clone(),
title: String::new(),
content: post.text.clone(),
media: post.media,
sensitive: post.sensitive,
site_id: "bsky",
@@ -410,7 +413,8 @@ mod tests {
fetched.source_url,
"https://bsky.app/profile/user.bsky.social/post/3xxxx"
);
assert_eq!(fetched.title, "hello <world>");
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "hello <world>");
assert_eq!(fetched.media.len(), 1);
assert!(!fetched.sensitive);
// display_name absent -> empty fallback
+17 -11
View File
@@ -123,21 +123,23 @@ impl From<model::Note> for Fetched {
let cw = content.cw.as_deref().unwrap_or_default();
// Notes carry hashtags inline in the text (no structured tags array);
// a CW note gets the marker prefixed so recipients see the spoiler.
let mut title = cw.to_string();
if !cw.is_empty() && !title.ends_with(' ') {
title.push(' ');
let mut text = cw.to_string();
if !cw.is_empty() && !text.ends_with(' ') {
text.push(' ');
}
title.push_str(content.text.as_deref().unwrap_or_default().trim());
let title = title.trim().to_string();
text.push_str(content.text.as_deref().unwrap_or_default().trim());
let text = text.trim().to_string();
let caption = caption(&url, &author_url, &author, &title);
let caption = caption(&url, &author_url, &author, &text);
let sensitive = content.cw.is_some() || content.files.iter().any(|f| f.is_sensitive);
let media: Vec<Media> = content.files.iter().filter_map(media_from_file).collect();
Fetched {
source_url: url.clone(),
caption,
title: title.clone(),
// A note has no title: its text (CW marker included) is content.
title: String::new(),
content: text.clone(),
media,
sensitive,
site_id: "misskey",
@@ -145,7 +147,8 @@ impl From<model::Note> for Fetched {
url,
author: encode_text(&author).into_owned(),
author_url: author_url.clone(),
title: encode_text(&title).into_owned(),
title: String::new(),
content: encode_text(&text).into_owned(),
tags: String::new(),
}),
_keep_alive: None,
@@ -257,7 +260,8 @@ mod tests {
"https://misskey.io/notes/aotihl10lqrs015s"
);
assert_eq!(fetched.site_id, "misskey");
assert_eq!(fetched.title, "hello");
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "hello");
assert!(fetched.sensitive);
assert_eq!(fetched.media.len(), 1);
match &fetched.media[0] {
@@ -308,7 +312,8 @@ mod tests {
note["text"] = serde_json::json!("body");
let fetched: Fetched = note_json(note).into();
assert!(fetched.sensitive);
assert_eq!(fetched.title, "spoiler body");
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "spoiler body");
}
#[test]
@@ -341,7 +346,8 @@ mod tests {
}
});
let fetched: Fetched = note_json(note).into();
assert_eq!(fetched.title, "inner text");
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "inner text");
assert_eq!(fetched.media.len(), 1);
// The source URL still points at the renote shell the user posted.
assert_eq!(
+161 -32
View File
@@ -1,8 +1,8 @@
//! Site fetching dispatcher and unified result types.
//!
//! Dispatch order: twitter → bsky → misskey → pixiv. Each site module
//! exports a `PATTERN`, `enabled()` and `fetch_from_url()`; a future site
//! plugs in by adding one guarded entry in `SITES`.
//! Dispatch order: twitter → bsky → misskey → pixiv → bilibili. Each site
//! module exports a `PATTERN`, `enabled()` and `fetch_from_url()`; a future
//! site plugs in by adding one guarded entry in `SITES`.
use std::future::Future;
use std::pin::Pin;
@@ -13,6 +13,7 @@ use std::time::Duration;
use regex::Regex;
use thiserror::Error;
pub mod bilibili;
pub mod bsky;
pub mod misskey;
pub mod pixiv;
@@ -20,22 +21,31 @@ pub mod twitter;
pub use pixiv::PixivError;
/// The result of fetching a post: canonical URL, HTML caption, raw text,
/// media list and spoiler flag. Produced by [`fetch`].
/// The result of fetching a post: canonical URL, HTML caption, the post's
/// title and body, media list and spoiler flag. Produced by [`fetch`].
#[derive(Debug)]
pub struct Fetched {
/// Canonical URL: `x.com/{author}/status/{id}` |
/// `https://www.pixiv.net/artworks/{id}` |
/// `https://bsky.app/profile/{handle}/post/{rkey}`
/// `https://bsky.app/profile/{handle}/post/{rkey}` |
/// `https://www.bilibili.com/opus/{id}`
pub source_url: String,
/// The exact HTML produced by the site's caption().
pub caption: String,
/// Raw post text (tweet text / bsky text / pixiv title).
/// The post's own title, where the platform has one: a pixiv artwork's
/// title, the headline of a bilibili opus post or the title of the video
/// an AV dynamic attaches. Empty on the platforms whose posts are text
/// only (x/twitter, bsky, misskey) and on bilibili posts without a
/// headline.
pub title: String,
/// The post's body text, as the platform exposes it: a tweet, a bsky or
/// misskey post, a bilibili dynamic's text, a pixiv artwork's description
/// (HTML flattened). Empty when the post has no text at all.
pub content: String,
pub media: Vec<crate::media::Media>,
/// Spoiler flag for all media of this post.
pub sensitive: bool,
/// Site id (`"twitter"` / `"bsky"` / `"pixiv"`): the single source of
/// Site id (`"twitter"` / `"bsky"` / `"pixiv"` / `"bilibili"`): the single source of
/// truth for site identity — caption-format lookup, cache-key prefix and
/// the SetFormat whitelist all derive from it. Set by the producing site.
pub site_id: &'static str,
@@ -46,24 +56,38 @@ pub struct Fetched {
pub(crate) _keep_alive: Option<tempfile::TempDir>,
}
/// Values for the `{url} {author} {author_url} {title} {tags}` placeholders in
/// user-supplied caption formats, substituted by [`caption_from_fields`] as
/// HTML text (never as an attribute value).
/// Values for the `{url} {author} {author_url} {title} {content} {tags}`
/// placeholders in user-supplied caption formats, substituted by
/// [`caption_from_fields`] as HTML text (never as an attribute value).
///
/// `author`, `title` and `tags` come from the site API (post text, display
/// names) and are HTML-escaped at construction. `url` and `author_url` stay
/// raw: they are canonical URLs the adapter builds from numeric ids and
/// API-constrained handles/DIDs, so they carry no escapable character — the
/// bot's `/test` report relies on that when it embeds them.
/// `author`, `title`, `content` and `tags` come from the site API (post
/// text, display names, descriptions) and are HTML-escaped at construction.
/// `url` and `author_url` stay raw: they are canonical URLs the adapter
/// builds from numeric ids and API-constrained handles/DIDs, so they carry
/// no escapable character — the bot's `/test` report relies on that when it
/// embeds them.
#[derive(Debug)]
pub(crate) struct RenderData {
pub url: String,
pub author: String,
pub author_url: String,
pub title: String,
pub content: String,
pub tags: String,
}
/// The post's text as one string: title and content joined by a line break,
/// each only when it is non-empty. This is what the sites' built-in captions
/// show after the author line, and what the bot quotes when it is long.
pub fn compose_text(title: &str, content: &str) -> String {
match (title.is_empty(), content.is_empty()) {
(false, false) => format!("{title}\n{content}"),
(false, true) => title.to_string(),
(true, false) => content.to_string(),
(true, true) => String::new(),
}
}
impl Fetched {
/// The site this post came from (used for per-site format overrides).
/// A thin alias over [`Fetched::site_id`] kept for callers that read the
@@ -87,21 +111,23 @@ impl Fetched {
&data.author,
&data.author_url,
&data.title,
&data.content,
&data.tags,
),
_ => truncate_caption(&self.caption),
}
}
/// The pre-escaped placeholder values (author, author_url, title, tags)
/// a caller needs to rebuild a caption later, e.g. for a cached post
/// where the [`Fetched`] is no longer available.
pub fn render_fields(&self) -> Option<(&str, &str, &str, &str)> {
/// The pre-escaped placeholder values (author, author_url, title,
/// content, tags) a caller needs to rebuild a caption later, e.g. for a
/// cached post where the [`Fetched`] is no longer available.
pub fn render_fields(&self) -> Option<(&str, &str, &str, &str, &str)> {
self.render_data.as_ref().map(|d| {
(
d.author.as_str(),
d.author_url.as_str(),
d.title.as_str(),
d.content.as_str(),
d.tags.as_str(),
)
})
@@ -147,6 +173,11 @@ pub fn truncate_caption(caption: &str) -> String {
/// [`Fetched::caption_with`]. An empty format returns `built_in` unchanged.
/// The result is truncated to [`MAX_CAPTION_CHARS`] (Telegram's caption
/// limit for HTML parse mode).
///
/// One flat argument per placeholder keeps the two callers (the fresh and the
/// cached caption path) mirroring each other; the same shape as the bot's
/// `debug_report`.
#[allow(clippy::too_many_arguments)]
pub fn caption_from_fields(
format: &str,
built_in: &str,
@@ -154,6 +185,7 @@ pub fn caption_from_fields(
author: &str,
author_url: &str,
title: &str,
content: &str,
tags: &str,
) -> String {
if format.is_empty() {
@@ -166,6 +198,7 @@ pub fn caption_from_fields(
.replace("{author}", author)
.replace("{author_url}", author_url)
.replace("{title}", title)
.replace("{content}", content)
.replace("{tags}", tags),
)
}
@@ -214,6 +247,12 @@ pub enum FetchError {
NotFound,
#[error("blocked")]
Blocked,
/// The URL matches a registered site that is disabled right now (pixiv
/// without `PIXIV_REFRESH_TOKEN`, or after a failed login). Distinct from
/// `Ok(None)` — an unsupported link — so the bot can tell the user why
/// the link was not handled instead of silently ignoring it.
#[error("{site} support is disabled")]
Disabled { site: &'static str },
/// The post exists but its content is withheld (twitter NSFW /
/// age-restricted tweets come back as an empty `{}` from syndication).
#[error("content withheld (sensitive)")]
@@ -283,7 +322,7 @@ pub(crate) fn log_once_ffmpeg_missing() {
}
/// Site adapter: one impl per supported site (twitter / bsky / misskey /
/// pixiv), registered in `SITES`. All site-specific knowledge — URL pattern,
/// pixiv / bilibili), registered in `SITES`. All site-specific knowledge — URL pattern,
/// cache-key format, fetch, retry policy, media-host headers, startup
/// validation — lives in the site module; the central dispatcher only
/// iterates the registry.
@@ -294,9 +333,9 @@ pub(crate) fn log_once_ffmpeg_missing() {
/// site structs are stateless unit structs, so the boxed futures never
/// borrow from `self` beyond the call's scope.
pub trait Site: Send + Sync {
/// Stable site id (`"twitter"` / `"bsky"` / `"misskey"` / `"pixiv"`):
/// caption-format lookup, cache-key prefixes and the SetFormat whitelist
/// derive from it.
/// Stable site id (`"twitter"` / `"bsky"` / `"misskey"` / `"pixiv"` /
/// `"bilibili"`): caption-format lookup, cache-key prefixes and the
/// SetFormat whitelist derive from it.
fn id(&self) -> &'static str;
/// URL pattern; the dispatcher's first match wins (dispatch order).
fn pattern(&self) -> &'static Regex;
@@ -332,14 +371,15 @@ pub trait Site: Send + Sync {
type SiteFuture<'a, T, E = FetchError> = Pin<Box<dyn Future<Output = Result<T, E>> + Send + 'a>>;
/// The one registry of supported sites, in dispatch order (twitter → bsky →
/// misskey → pixiv). Adding a site = new module + one `Box::new(...)` entry
/// here; the bot crate never lists sites itself.
/// misskey → pixiv → bilibili). Adding a site = new module + one
/// `Box::new(...)` entry here; the bot crate never lists sites itself.
static SITES: LazyLock<Vec<Box<dyn Site>>> = LazyLock::new(|| {
vec![
Box::new(twitter::TwitterSite),
Box::new(bsky::BskySite),
Box::new(misskey::MisskeySite),
Box::new(pixiv::PixivSite),
Box::new(bilibili::BilibiliSite),
]
});
@@ -351,6 +391,17 @@ fn find_site(url: &str) -> Option<&'static dyn Site> {
.map(|site| site.as_ref())
}
/// The site whose pattern matches `url` but which is disabled right now.
/// `None` when no site matches the URL at all, or when the matching site is
/// enabled. Lets the dispatcher tell "unsupported link" (silently ignored)
/// apart from "this bot has that site switched off" (reported to the user).
fn disabled_site(url: &str) -> Option<&'static str> {
SITES
.iter()
.find(|site| !site.enabled() && site.pattern().is_match(url))
.map(|site| site.id())
}
/// Every supported site id, in dispatch order. The bot's SetFormat whitelist
/// derives from this list.
pub fn site_ids() -> Vec<&'static str> {
@@ -374,7 +425,9 @@ pub async fn validate_all() -> Vec<(&'static str, String)> {
}
/// Fetches a post from its URL. Returns `Ok(None)` when no site pattern
/// matches (unsupported links are silently ignored by the bot).
/// matches (unsupported links are silently ignored by the bot) and
/// [`FetchError::Disabled`] when the URL belongs to a registered site that is
/// switched off right now — the two are different answers for the user.
///
/// Transient failures are retried: 3 total attempts with 1s then 2s delays.
/// What counts as transient is the matched site's own policy (`is_retryable`
@@ -398,7 +451,13 @@ const MAX_FETCH_ATTEMPTS: u32 = 3;
async fn fetch_with_attempts(url: &str, attempts: u32) -> Result<Option<Fetched>, FetchError> {
let Some(site) = find_site(url) else {
return Ok(None);
// A registered-but-disabled site (pixiv without a token) is not an
// unsupported link: report it, so the bot answers the user instead of
// ignoring the message.
return match disabled_site(url) {
Some(site) => Err(FetchError::Disabled { site }),
None => Ok(None),
};
};
for attempt in 0..attempts.max(1) {
match site.fetch_from_url(url).await {
@@ -424,6 +483,15 @@ async fn fetch_with_attempts(url: &str, attempts: u32) -> Result<Option<Fetched>
unreachable!("retry loop always returns")
}
/// Whether fetching `url` requires site-specific headers (pixiv's `Referer`
/// for `pximg.net` hotlink protection, see [`Site::media_headers`]). Telegram's
/// own fetch of a media URL sends none of them, so a URL that needs them fails
/// there — callers that hand a URL to Telegram (inline query results) must
/// skip such media instead of shipping a broken item.
pub fn needs_media_headers(url: &str) -> bool {
SITES.iter().any(|site| site.media_headers(url).is_some())
}
/// Applies every site's media-header rule to a download request (pixiv's
/// `Referer` for pximg.net hotlink protection). Sites contribute via their
/// `media_headers(url)` — the central download code carries no per-site logic.
@@ -541,6 +609,10 @@ mod tests {
cache_key("https://bsky.app/profile/handle.example/post/3lorem"),
Some("bsky:handle.example/3lorem".into())
);
assert_eq!(
cache_key("https://t.bilibili.com/1245284537985925159"),
Some("bilibili:1245284537985925159".into())
);
assert_eq!(cache_key("https://example.com/not-a-post"), None);
}
@@ -549,16 +621,21 @@ mod tests {
assert_eq!(site_id_from_key("twitter:123"), "twitter");
assert_eq!(site_id_from_key("pixiv:123"), "pixiv");
assert_eq!(site_id_from_key("bsky:handle.example/3lorem"), "bsky");
assert_eq!(site_id_from_key("bilibili:123"), "bilibili");
assert_eq!(site_id_from_key("unknown:1"), "unknown");
assert_eq!(site_id_from_key("no-colon"), "unknown");
}
#[test]
fn registry_lists_all_sites_in_dispatch_order() {
assert_eq!(site_ids(), vec!["twitter", "bsky", "misskey", "pixiv"]);
assert_eq!(
site_ids(),
vec!["twitter", "bsky", "misskey", "pixiv", "bilibili"]
);
// Enabled sites dispatch; unsupported URLs never match.
assert!(find_site("https://x.com/u/status/1").is_some());
assert!(find_site("https://misskey.io/notes/abc").is_some());
assert!(find_site("https://t.bilibili.com/1245284537985925159").is_some());
assert!(find_site("https://example.com/x").is_none());
// Cache keys are pattern-driven, independent of the enabled() gate
// (pixiv is disabled in tests without PIXIV_REFRESH_TOKEN).
@@ -586,25 +663,34 @@ mod tests {
// The format string is escaped, the field values are substituted
// verbatim (callers pass the already-escaped render data).
let out = caption_from_fields(
"see {author} at {url} — {title}",
"see {author} at {url} — {title}: {content}",
"",
"https://x.com/u/status/1",
"A &amp; B",
"https://x.com/u",
"hello <world>",
"the body",
"",
);
assert_eq!(
out,
"see A &amp; B at https://x.com/u/status/1 — hello <world>"
"see A &amp; B at https://x.com/u/status/1 — hello <world>: the body"
);
// Empty format keeps the built-in caption untouched.
assert_eq!(
caption_from_fields("", "built-in", "u", "a", "au", "t", "g"),
caption_from_fields("", "built-in", "u", "a", "au", "t", "c", "g"),
"built-in"
);
}
#[test]
fn compose_text_joins_title_and_content() {
assert_eq!(compose_text("标题", "正文"), "标题\n正文");
assert_eq!(compose_text("标题", ""), "标题");
assert_eq!(compose_text("", "正文"), "正文");
assert_eq!(compose_text("", ""), "");
}
#[test]
fn truncate_caption_keeps_short_text() {
assert_eq!(truncate_caption("short"), "short");
@@ -656,6 +742,49 @@ mod tests {
assert!(matches!(result, Ok(None)), "got {result:?}");
}
#[test]
fn media_headers_are_reported_only_where_telegram_would_fail() {
// pixiv's CDN needs a Referer, which only the bot can send: an inline
// result pointing at it renders broken, so callers skip it.
assert!(needs_media_headers(
"https://i.pximg.net/img-original/img/2024/01/01/00/00/00/1_p0.jpg"
));
// The rest serve direct requests (verified per site in their modules).
for url in [
"https://pbs.twimg.com/media/1.jpg",
"https://cdn.bsky.app/img/1.jpg",
"https://media.misskeyusercontent.jp/io/1.webp",
"https://i0.hdslb.com/bfs/1.jpg",
] {
assert!(!needs_media_headers(url), "{url}");
}
}
#[tokio::test]
async fn disabled_site_is_reported_not_ignored() {
// pixiv is the only token-gated site; with PIXIV_REFRESH_TOKEN set it
// is enabled and this link would hit the network, so skip then.
if std::env::var("PIXIV_REFRESH_TOKEN")
.ok()
.filter(|s| !s.is_empty())
.is_some()
{
eprintln!("skipping: PIXIV_REFRESH_TOKEN is set");
return;
}
let result = fetch("https://www.pixiv.net/artworks/1").await;
assert!(
matches!(result, Err(FetchError::Disabled { site: "pixiv" })),
"got {result:?}"
);
// The cache key still resolves: the bot keys the reply and the link
// cache off it even when the site is off.
assert_eq!(
cache_key("https://www.pixiv.net/artworks/1"),
Some("pixiv:1".into())
);
}
#[tokio::test]
async fn download_media_pixiv_original_with_referer() {
// Proves the Referer header is attached for i.pximg.net: a header-less
@@ -107,10 +107,60 @@ pub fn media_headers(url: &str) -> Option<Vec<(&'static str, String)>> {
}
}
/// Flattens the app API's HTML description into plain text: `<br>` (and `<p>`)
/// become line breaks, other tags are dropped, entities decoded, the ends
/// trimmed. A caption shows text, not markup, so the author's `<a href>` links
/// contribute their link text only.
fn flatten_html(raw: &str) -> String {
let mut out = String::with_capacity(raw.len());
let mut chars = raw.chars().peekable();
while let Some(c) = chars.next() {
// Only `<` followed by `/` or a letter opens a tag — a bare `<` in
// prose ("2 < 3") is text.
let opens_tag = c == '<'
&& chars
.peek()
.is_some_and(|next| *next == '/' || next.is_ascii_alphabetic());
if !opens_tag {
out.push(c);
continue;
}
let mut tag = String::new();
let mut closed = false;
for c in chars.by_ref() {
if c == '>' {
closed = true;
break;
}
tag.push(c);
}
if !closed {
// Unclosed `<…`: keep it as text rather than dropping the tail.
out.push('<');
out.push_str(&tag);
break;
}
// `<br>`, `<br/>`, `<br />` with or without attributes, and both
// halves of a paragraph break the line; everything else is dropped.
let tag = tag
.trim()
.trim_start_matches('/')
.trim_end_matches('/')
.trim()
.to_ascii_lowercase();
if tag == "p" || tag.starts_with("br") {
out.push('\n');
}
}
html_escape::decode_html_entities(&out).trim().to_string()
}
#[derive(Debug)]
pub struct Illustration {
id: String,
title: String,
/// The artwork's description, HTML flattened to plain text.
content: String,
author: String,
author_id: String,
tags: Vec<String>,
@@ -150,6 +200,7 @@ impl Illustration {
pub fn from_model(model: &IllustrationModel) -> Self {
let id = model.id.to_string();
let title = model.title.clone();
let content = flatten_html(&model.caption);
let author = model.user.name.clone();
let author_id = model.user.id.to_string();
let mut tags: Vec<String> = model.tags.iter().map(|tag| tag.name.clone()).collect();
@@ -193,6 +244,7 @@ impl Illustration {
Self {
id,
title,
content,
author,
author_id,
tags,
@@ -218,12 +270,14 @@ impl From<Illustration> for Fetched {
author: encode_text(&illustration.author).into_owned(),
author_url: author_url.clone(),
title: encode_text(&illustration.title).into_owned(),
content: encode_text(&illustration.content).into_owned(),
tags: encode_text(&tags).into_owned(),
});
Fetched {
source_url: url,
caption: illustration.caption(),
title: illustration.title.clone(),
content: illustration.content.clone(),
media: illustration.media,
sensitive: illustration.nsfw,
site_id: "pixiv",
@@ -262,6 +316,7 @@ mod tests {
"illust": {
"id": 123,
"title": "Art <title>",
"caption": "一行说明<br />二行 <a href=\"https://x.example/\">链接</a> &amp; 结尾",
"type": type_,
"image_urls": {
"medium": "medium.jpg",
@@ -284,6 +339,44 @@ mod tests {
Illustration::from_model(&model)
}
/// The description arrives as HTML and becomes plain-text content: breaks
/// kept, tags dropped (links keep their text), entities decoded.
#[test]
fn from_json_maps_description_to_content() {
let v = illust_json("illust", 1, None, Some("o.jpg"), vec![], 0);
let illustration = parse(v);
assert_eq!(illustration.content, "一行说明\n二行 链接 & 结尾");
let fetched: Fetched = illustration.into();
assert_eq!(fetched.title, "Art <title>");
assert_eq!(fetched.content, "一行说明\n二行 链接 & 结尾");
// The built-in caption keeps its layout: the description stays out of
// it and is available through `{content}`.
assert!(!fetched.caption.contains("一行说明"), "{}", fetched.caption);
assert_eq!(
fetched.render_fields().unwrap().3,
"一行说明\n二行 链接 &amp; 结尾"
);
assert!(
fetched
.caption_with("{title}: {content}")
.ends_with("一行说明\n二行 链接 &amp; 结尾")
);
}
#[test]
fn flatten_html_handles_common_markup() {
assert_eq!(flatten_html(""), "");
assert_eq!(flatten_html("plain"), "plain");
assert_eq!(flatten_html("a<br />b<br/>c<br>d"), "a\nb\nc\nd");
// A paragraph break is a blank line, exactly like `<br /><br />` —
// writing it as one newline would flatten the author's paragraphs.
assert_eq!(flatten_html("<p>one</p><p>two</p>"), "one\n\ntwo");
assert_eq!(flatten_html("a &amp; b &lt;c&gt;"), "a & b <c>");
// Nothing to strip: angle brackets that are not a tag survive.
assert_eq!(flatten_html("2 < 3"), "2 < 3");
}
#[test]
fn pattern_matches_all_forms() {
let cases = [
+4
View File
@@ -6,6 +6,10 @@ use serde::Deserialize;
pub struct IllustrationModel {
pub id: u64,
pub title: String,
/// The artwork's description as the app API returns it — HTML in most
/// works (`<br />`, `<a href>`, sometimes `<p>`), empty for many.
#[serde(default)]
pub caption: String,
pub r#type: TypeModel,
pub image_urls: ImageUrlsModel,
pub user: UserInfoModel,
+2 -1
View File
@@ -327,7 +327,8 @@ mod tests {
"https://x.com/nsfw_author/status/2083868672721039569"
);
// The appended media short link (no URL-entity mapping) is stripped.
assert_eq!(fetched.title, "nsfw content");
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "nsfw content");
}
#[test]
+18 -32
View File
@@ -43,27 +43,25 @@ pub async fn fetch_from_url(url: &str) -> Result<Fetched, FetchError> {
match fetch(id).await {
Ok(tweet) => Ok(tweet.into()),
// Syndication withholds NSFW/age-restricted tweets (empty `{}`).
// Retry as the logged-in user when TWITTER_AUTH_TOKEN is set;
// otherwise degrade to an empty result (the bot replies
// "No media found").
// Retry as the logged-in user when TWITTER_AUTH_TOKEN is set; without
// the token the withholding is reported as `Sensitive`, so the bot can
// answer "age-restricted / needs TWITTER_AUTH_TOKEN" instead of the
// misleading "No media found".
Err(FetchError::Sensitive) => {
if super::auth::enabled() {
match super::auth::fetch(id).await {
Ok(tweet) => Ok(tweet.into()),
// The tweet is genuinely gone (deleted / suspended /
// tombstoned): report it instead of degrading to an
// empty result ("No media found"). Only unexpected
// fallback failures (network, parse) keep the NSFW
// placeholder.
Err(FetchError::NotFound) => Err(FetchError::NotFound),
// Deleted/suspended (tombstoned) and unexpected fallback
// failures keep their own class: the bot reports what
// actually happened rather than "No media found".
Err(e) => {
log::warn!("twitter auth fallback failed for {id}: {e}");
Ok(empty_fetched(url))
Err(e)
}
}
} else {
log::debug!("tweet {id} is sensitive; set TWITTER_AUTH_TOKEN to fetch NSFW media");
Ok(empty_fetched(url))
Err(FetchError::Sensitive)
}
}
Err(e) => Err(e),
@@ -90,23 +88,6 @@ pub fn media_headers(_url: &str) -> Option<Vec<(&'static str, String)>> {
None
}
/// A Fetched with no media for withheld tweets: the bot replies
/// "No media found" and moves on instead of erroring.
fn empty_fetched(url: &str) -> Fetched {
Fetched {
source_url: url.to_string(),
// The raw user-supplied URL goes into an HTML caption; escape it so
// crafted links cannot break the parse (Telegram 400).
caption: encode_text(url).into_owned(),
title: String::new(),
media: vec![],
sensitive: true,
site_id: "twitter",
render_data: None,
_keep_alive: None,
}
}
/// Fetches a tweet from the syndication endpoint. Deleted/blocked tweets
/// surface as `FetchError::NotFound`; withheld content (empty tombstone,
/// age-restricted) as `FetchError::Sensitive`.
@@ -352,17 +333,20 @@ impl From<Tweet> for Fetched {
fn from(tweet: Tweet) -> Self {
let url = tweet.url();
let author_url = tweet.author_url();
// A tweet has no title: its text is all content.
let render_data = Some(crate::site::RenderData {
url: url.clone(),
author: encode_text(&tweet.author).into_owned(),
author_url: author_url.clone(),
title: encode_text(&tweet.text).into_owned(),
title: String::new(),
content: encode_text(&tweet.text).into_owned(),
tags: String::new(),
});
Fetched {
source_url: url,
caption: tweet.caption(),
title: tweet.text.clone(),
title: String::new(),
content: tweet.text.clone(),
media: tweet.media,
sensitive: tweet.sensitive,
site_id: "twitter",
@@ -437,7 +421,8 @@ mod tests {
assert_eq!(tweet.text, ">^ω^< & more 'quoted'");
assert_eq!(tweet.author, "O'Brien");
let fetched: Fetched = tweet.into();
assert_eq!(fetched.title, ">^ω^< & more 'quoted'");
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, ">^ω^< & more 'quoted'");
// The caption escapes the raw text exactly once (encode_text covers
// & < >; apostrophes stay literal — they are harmless in text).
assert!(
@@ -496,7 +481,8 @@ mod tests {
fetched.source_url,
"https://x.com/author_handle/status/861627479294746624"
);
assert_eq!(fetched.title, "a & b <c>");
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "a & b <c>");
assert!(fetched.sensitive);
assert_eq!(fetched.media.len(), 2);
match &fetched.media[0] {
+3 -3
View File
@@ -1,6 +1,6 @@
[package]
name = "xmedia-bot"
version = "1.6.0"
version = "1.8.0"
edition = "2024"
[dependencies]
@@ -13,8 +13,8 @@ pretty_env_logger = "0.5"
dotenv = "0.15"
url = "2.5.2"
html-escape = "0.2"
rusqlite = { version = "0.32", features = ["bundled"] }
rand = "0.8"
rusqlite = { version = "0.40", features = ["bundled"] }
rand = "0.10"
tempfile = "3"
parking_lot = "0.12"
bytes = "1"
+8 -1
View File
@@ -1,5 +1,6 @@
//! Central env handling. The only other places that read env are
//! `Bot::from_env` (TELOXIDE_TOKEN) and x-media (PIXIV_REFRESH_TOKEN).
//! `Bot::from_env` (TELOXIDE_TOKEN) and x-media (PIXIV_REFRESH_TOKEN,
//! TWITTER_AUTH_TOKEN, BILIBILI_COOKIE).
use std::env;
use std::net::IpAddr;
@@ -12,6 +13,10 @@ pub struct Config {
pub edit_message_ttl: Duration,
/// LINK_CACHE_TTL_SECONDS, default 604800 (7 days).
pub link_cache_ttl: Duration,
/// CAPTION_QUOTE_TEXT_CHARS, default 200: a post whose text (title plus
/// content) is at least this many characters gets that text wrapped in an
/// expandable blockquote inside its caption. `0` disables the wrap.
pub caption_quote_text_chars: usize,
// Webhook settings (moved out of main; names/defaults unchanged).
pub webhook_enabled: bool,
pub webhook_url: Option<url::Url>,
@@ -57,6 +62,7 @@ impl Config {
Duration::from_secs(parse_u64("EDIT_MESSAGE_TTL_SECONDS", 24 * 3600));
let link_cache_ttl =
Duration::from_secs(parse_u64("LINK_CACHE_TTL_SECONDS", 7 * 24 * 3600));
let caption_quote_text_chars = parse_u64("CAPTION_QUOTE_TEXT_CHARS", 200) as usize;
let webhook_enabled = env::var("WEBHOOK")
.is_ok_and(|v| matches!(v.to_lowercase().as_str(), "true" | "yes" | "1"));
@@ -92,6 +98,7 @@ impl Config {
admin_ids,
edit_message_ttl,
link_cache_ttl,
caption_quote_text_chars,
webhook_enabled,
webhook_url,
webhook_listen,
+6
View File
@@ -88,6 +88,12 @@ pub(crate) mod test_support {
&self.chat_store
}
/// The parsed config, mutable so a test can pin a knob (e.g. the
/// caption-quote threshold) instead of depending on the environment.
pub(crate) fn config_mut(&mut self) -> &mut Config {
&mut self.config
}
pub(crate) fn link_cache(&self) -> &LinkCache {
&self.link_cache
}
@@ -14,6 +14,8 @@ use teloxide::types::{CallbackQuery, CallbackQueryId, MessageId};
/// The `"forward"` button's data.
const FORWARD: &str = "forward";
/// The `"skip"` button's data: drop the prompt without forwarding.
const SKIP: &str = "skip";
/// Prefix of a template button's data: `"template|<name>"`.
const TEMPLATE_PREFIX: &str = "template|";
@@ -70,6 +72,29 @@ async fn handle_callback(
}
log::info!("callback from {chat_id} on prompt {prompt_message_id}: {data}");
if data == SKIP {
// Skip works with or without a forward channel: it is the explicit
// "do not forward this" answer, and it drops the record so the forward
// can never happen later.
log::info!("edit-before-forward prompt {prompt_message_id} skipped");
ctx.chat_store
.update(chat_id, |data| {
data.edit_message.remove(&prompt_message_id);
})
.await;
let _ = ctx
.sender
.delete_message(ChatId(chat_id), MessageId(prompt_message_id as i32))
.await;
let _ = ctx
.sender
.answer_callback_query(
callback_query_id,
Some("Skipped — nothing was forwarded.".to_string()),
)
.await;
return;
}
if data == FORWARD {
match chat_data.forward_channel_id {
Some(channel_id) => {
@@ -244,6 +269,31 @@ mod tests {
);
}
#[tokio::test]
async fn skip_drops_the_prompt_without_forwarding() {
// "skip" needs no forward channel and no scripted outcomes: it deletes
// the prompt and drops the record, so no forward can ever happen.
let sender = MockSender::scripted(vec![], api_error);
let stores = TestStores::new();
let ctx = stores.ctx(&sender);
seed_prompt(&ctx, crate::db::unix_now()).await;
handle_callback(&ctx, callback_id(), 1, PROMPT_ID, "skip").await;
assert_eq!(
sender.calls(),
vec!["delete_message", "answer_callback_query"]
);
assert_eq!(
sender.answers(),
vec![Some("Skipped — nothing was forwarded.".to_string())]
);
assert!(
ctx.chat_store.get(1).await.edit_message.is_empty(),
"a skipped prompt must drop its record"
);
}
#[tokio::test]
async fn forward_without_a_channel_is_reported() {
let sender = MockSender::scripted(vec![], api_error);
+558 -18
View File
@@ -4,6 +4,7 @@
use super::urls::{PostSend, url_media};
use super::{CHAT_STORE, CONFIG, LINK_CACHE, log_key, reply, reply_html};
use crate::ctx::AppContext;
use crate::state::ChatData;
use teloxide::RequestError;
use teloxide::prelude::*;
use teloxide::types::{ChatId, Message, Recipient};
@@ -33,13 +34,20 @@ pub(crate) enum Command {
parse_with = "split"
)]
SetTemplate(String),
#[command(description = "Remove a saved template", parse_with = "split")]
RemoveTemplate(String),
#[command(description = "Show this chat's settings")]
Settings,
#[command(description = "Show chat state (debug; admin only)")]
BotDict,
#[command(description = "Set site caption format", parse_with = "split")]
#[command(
description = "Set site caption format (- to reset)",
parse_with = parse_arg_remainder
)]
SetFormat(String),
#[command(
description = "Clear link cache (admin; optional URL, else all)",
parse_with = "split"
parse_with = parse_arg_remainder
)]
ClearCache(String),
#[command(
@@ -62,6 +70,118 @@ fn parse_arg_remainder(s: String) -> Result<(String,), ParseError> {
Ok((s.trim().to_string(),))
}
/// Placeholders `/set_format` accepts, mirroring what
/// `x_media::site::caption_from_fields` substitutes.
const FORMAT_PLACEHOLDERS: [&str; 6] = ["url", "author", "author_url", "title", "content", "tags"];
/// `/start`'s welcome: what the bot is for, where links work, where to look
/// next. The old "Hello!" left a first-time user with nothing.
const START_TEXT: &str = "\
Send me a post link and I'll send back its images, videos and GIFs with the title, author and tags.
Supported: X/Twitter, Pixiv, Bluesky, Misskey (misskey.io), Bilibili.
In a private chat just paste the link. In a group, use inline mode (type @, pick me, then the link).
/help lists every command.";
/// Appended to `/help`'s command list: argument syntax, caption
/// placeholders and the private-chat rule — none of which teloxide's
/// `descriptions()` renders (it prints `/command — description` only).
const HELP_FOOTER: &str = "\
Arguments
/set_forward_channel <@channel or channel id>
/set_template <name> — reply to a message containing [] to save it
/remove_template <name> — see /settings for the saved names
/set_format <site> <format> — '-' restores the built-in format
/test <link> / /debug <link>
Caption placeholders (for /set_format)
{url} {author} {author_url} {title} {content} {tags}
A template's [] is replaced by the post link when forwarding.
Links are handled in private chats only; in a group use inline mode.";
/// Cap on template names echoed by `/settings`: a chat with hundreds of
/// templates must not produce a message Telegram rejects for length.
const MAX_SETTINGS_TEMPLATE_NAMES: usize = 30;
/// Sorted template names: the order `/settings`, `/remove_template` and the
/// prompt's buttons all show.
fn sorted_template_names(data: &ChatData) -> Vec<String> {
let mut names: Vec<String> = data.template.keys().cloned().collect();
names.sort();
names
}
/// `/settings`: what this chat is configured to do, readable by anyone in it
/// (unlike `/bot_dict`, which dumps the raw state and is admin-only).
fn settings_text(data: &ChatData) -> String {
let mut lines = Vec::new();
match data.forward_channel_id {
Some(id) => lines.push(format!("Forward channel: {id}")),
None => lines.push(
"Forward channel: not set (use /set_forward_channel <@channel or id>)".to_string(),
),
}
lines.push(format!(
"Edit before forward: {}",
if data.edit_before_forward {
"on"
} else {
"off"
}
));
let mut formats: Vec<String> = data
.message_format
.iter()
.map(|(site, format)| format!("{site} => {format}"))
.collect();
formats.sort();
lines.push(if formats.is_empty() {
"Caption formats: built-in for every site".to_string()
} else {
format!("Caption formats:\n {}", formats.join("\n "))
});
let names = sorted_template_names(data);
lines.push(match names.len() {
0 => "Templates: none".to_string(),
n => format!(
"Templates ({n}): {}{}",
names
.iter()
.take(MAX_SETTINGS_TEMPLATE_NAMES)
.cloned()
.collect::<Vec<_>>()
.join(", "),
if n > MAX_SETTINGS_TEMPLATE_NAMES {
format!(", +{} more", n - MAX_SETTINGS_TEMPLATE_NAMES)
} else {
String::new()
}
),
});
lines.join("\n")
}
/// The first `{…}` token in a caption format that is not a known placeholder
/// (`None` when all of them are). The renderer replaces exact keys only, so an
/// unknown token would be published verbatim in every caption of that site —
/// caught here instead.
fn unknown_placeholder(format: &str) -> Option<&str> {
let mut rest = format;
while let Some(open) = rest.find('{') {
let after = &rest[open + 1..];
// An unclosed `{` is not a placeholder token at all.
let close = after.find('}')?;
let token = &after[..close];
if !FORMAT_PLACEHOLDERS.contains(&token) {
return Some(token);
}
rest = &after[close + 1..];
}
None
}
enum SetForwardChannelError {
EmptyParameter,
NotChannel,
@@ -140,11 +260,17 @@ pub(crate) async fn execute_command(
) -> Result<(), RequestError> {
match command {
Command::Start => {
bot.send_message(message.chat.id, "Hello!").await?;
bot.send_message(message.chat.id, START_TEXT).await?;
}
Command::Help => {
bot.send_message(message.chat.id, Command::descriptions().to_string())
.await?;
// The command list plus the parts teloxide's `descriptions()`
// cannot show: argument syntax, caption placeholders, and where a
// link actually works.
bot.send_message(
message.chat.id,
format!("{}\n\n{}", Command::descriptions(), HELP_FOOTER),
)
.await?;
}
Command::SetForwardChannel(channel) => {
let result = match set_forward_channel_handler(bot, message, channel).await {
@@ -232,6 +358,41 @@ pub(crate) async fn execute_command(
};
reply(bot, message.chat.id.0, message.id, text).await?;
}
Command::RemoveTemplate(name) => {
let chat_id = message.chat.id.0;
let name = name.trim().to_string();
if name.is_empty() {
reply(
bot,
chat_id,
message.id,
"Usage: /remove_template <name> (see /settings for the saved names)",
)
.await?;
return Ok(());
}
let removed = CHAT_STORE
.update(chat_id, |data| data.template.remove(&name).is_some())
.await;
let text = if removed {
format!("Template '{name}' removed.")
} else {
// Name the live templates: a typo would otherwise look like a
// successful delete.
let names = sorted_template_names(&CHAT_STORE.get(chat_id).await);
if names.is_empty() {
format!("No template named '{name}'. None are saved yet.")
} else {
format!("No template named '{name}'. Saved: {}", names.join(", "))
}
};
reply(bot, chat_id, message.id, text).await?;
}
Command::Settings => {
let chat_id = message.chat.id.0;
let data = CHAT_STORE.get(chat_id).await;
reply(bot, chat_id, message.id, settings_text(&data)).await?;
}
Command::BotDict => {
// Debug dump of the chat's persisted state: admin only (it echoes
// forward-channel ids and templates to whoever asks).
@@ -279,7 +440,45 @@ pub(crate) async fn execute_command(
bot,
message.chat.id.0,
message.id,
"Unknown site. Use twitter, bsky, pixiv or misskey.",
"Unknown site. Use twitter, bsky, pixiv, misskey or bilibili.",
)
.await?;
return Ok(());
}
// `-` resets to the site's built-in caption: without it a chat that
// set a format once could never get back to the default (the
// built-in format string is not something a user can retype).
if format == "-" {
CHAT_STORE
.update(chat_id, |data| {
data.message_format.remove(site);
})
.await;
reply(
bot,
message.chat.id.0,
message.id,
"Format reset to the built-in one.",
)
.await?;
return Ok(());
}
// A typo like {titel} would otherwise be rendered literally into
// every caption of that site (the renderer only substitutes the
// exact keys), which is invisible until a post arrives.
if let Some(token) = unknown_placeholder(&format) {
reply(
bot,
message.chat.id.0,
message.id,
format!(
"Unknown placeholder {{{token}}}. Available: {}",
FORMAT_PLACEHOLDERS
.iter()
.map(|name| format!("{{{name}}}"))
.collect::<Vec<_>>()
.join(" ")
),
)
.await?;
return Ok(());
@@ -289,7 +488,13 @@ pub(crate) async fn execute_command(
data.message_format.insert(site.to_string(), format);
})
.await;
reply(bot, message.chat.id.0, message.id, "Format set.").await?;
reply(
bot,
message.chat.id.0,
message.id,
"Format set. Use /debug <link> to preview the caption.",
)
.await?;
}
Command::ClearCache(arg) => {
let sender_id = message
@@ -320,7 +525,7 @@ pub(crate) async fn execute_command(
bot,
message.chat.id.0,
message.id,
"Unrecognized link. Use a twitter/x, pixiv, bsky or misskey post URL.",
"Unrecognized link. Use a twitter/x, pixiv, bsky, misskey or bilibili post URL.",
)
.await?;
return Ok(());
@@ -358,7 +563,7 @@ pub(crate) async fn execute_command(
bot,
message.chat.id.0,
message.id,
"No enabled site matches this link (twitter/x, pixiv, bsky or misskey).",
"No enabled site matches this link (twitter/x, pixiv, bsky, misskey or bilibili).",
)
.await?;
return Ok(());
@@ -400,7 +605,7 @@ pub(crate) async fn execute_command(
bot,
message.chat.id.0,
message.id,
"No enabled site matches this link (twitter/x, pixiv, bsky or misskey).",
"No enabled site matches this link (twitter/x, pixiv, bsky, misskey or bilibili).",
)
.await?;
}
@@ -414,14 +619,33 @@ pub(crate) async fn execute_command(
.await?;
}
Ok(Some(fetched)) => {
// The preview must show what a link would actually send:
// the chat's per-site format override plus the long-post
// quoting. Rendering the raw built-in caption here made
// `/set_format` look like it did nothing.
let format = CHAT_STORE
.get(message.chat.id.0)
.await
.message_format
.get(fetched.site_name())
.cloned()
.unwrap_or_default();
let caption = preview_caption(
&format,
&fetched.caption,
&fetched.source_url,
fetched.render_fields(),
CONFIG.caption_quote_text_chars,
);
let report = debug_report(
url,
fetched.site_name(),
&fetched.source_url,
&fetched.title,
&fetched.content,
fetched.render_fields(),
fetched.sensitive,
&fetched.caption,
&caption,
&fetched.media,
);
// HTML report: the caption renders inside a <blockquote>
@@ -439,12 +663,32 @@ fn plural(n: usize) -> &'static str {
if n == 1 { "y" } else { "ies" }
}
/// Bot profile texts (Bot API `setMyDescription` / `setMyShortDescription`):
/// shown on the bot's profile page and in the share sheet. Without them a
/// shared link says nothing about what the bot does.
const BOT_DESCRIPTION: &str = "\
Send a post link from X/Twitter, Pixiv, Bluesky, Misskey (misskey.io) or Bilibili and get its images, videos and GIFs back with the title, author and tags.
Links are handled in private chats; a group can use inline mode. /help lists every command.";
const BOT_SHORT_DESCRIPTION: &str =
"Post links (X, Pixiv, Bluesky, Misskey, Bilibili) -> media messages";
/// Registers the bot's command list with Telegram so clients show it in the
/// `/` menu (Bot API `setMyCommands`).
/// `/` menu (Bot API `setMyCommands`), plus its profile description texts.
pub async fn register_commands(bot: &Bot) -> Result<(), RequestError> {
let commands = Command::bot_commands();
bot.set_my_commands(commands.clone()).await?;
log::info!("registered {} commands", commands.len());
// Profile texts are cosmetic: a failure (rare) must not abort startup.
if let Err(e) = bot.set_my_description().description(BOT_DESCRIPTION).await {
log::warn!("failed to set the bot description: {e}");
}
if let Err(e) = bot
.set_my_short_description()
.short_description(BOT_SHORT_DESCRIPTION)
.await
{
log::warn!("failed to set the bot short description: {e}");
}
Ok(())
}
@@ -456,6 +700,31 @@ const MAX_DEBUG_REPORT_CHARS: usize = 4000;
/// message, so it must stay under Telegram's 4096-char limit.
const MAX_DEBUG_DUMP_CHARS: usize = 3500;
/// The caption a link would actually send for this chat: the per-site format
/// override (empty = the site's built-in caption) and, on a long post, the
/// same text quoting the send paths apply. `/debug` shows this so the preview
/// cannot drift from what the send paths produce.
fn preview_caption(
format: &str,
built_in: &str,
url: &str,
fields: Option<(&str, &str, &str, &str, &str)>,
quote_chars: usize,
) -> String {
let caption = match fields {
// Same call the send paths make through `Fetched::caption_with`: an
// empty format falls back to the built-in caption.
Some((author, author_url, title, content, tags)) => x_media::site::caption_from_fields(
format, built_in, url, author, author_url, title, content, tags,
),
None => x_media::site::truncate_caption(built_in),
};
let text = fields
.map(|(_, _, title, content, _)| x_media::site::compose_text(title, content))
.unwrap_or_default();
crate::send::quote_long_caption(&caption, &text, quote_chars).into_owned()
}
/// Builds the HTML report for the `/debug` command: what the parser produced
/// for a link (site, canonical URL, title/author/tags, caption and the media
/// list) — no media is sent and nothing is cached or forwarded. Sent with
@@ -471,7 +740,8 @@ fn debug_report(
site_id: &str,
source_url: &str,
title: &str,
render: Option<(&str, &str, &str, &str)>,
content: &str,
render: Option<(&str, &str, &str, &str, &str)>,
sensitive: bool,
caption: &str,
media: &[x_media::media::Media],
@@ -491,7 +761,8 @@ fn debug_report(
html_escape::encode_text(source_url)
));
lines.push(format!("title: {}", html_escape::encode_text(title)));
if let Some((author, author_url, _title, tags)) = render {
lines.push(format!("content: {}", html_escape::encode_text(content)));
if let Some((author, author_url, _title, _content, tags)) = render {
// The render fields are already pre-escaped for HTML captions; embed
// them as-is so the report renders them exactly like the final
// caption. `author_url` is raw and gets escaped here.
@@ -536,7 +807,9 @@ fn debug_report(
#[cfg(test)]
mod tests {
use super::{MAX_DEBUG_REPORT_CHARS, debug_report};
use super::{
MAX_DEBUG_REPORT_CHARS, debug_report, preview_caption, settings_text, unknown_placeholder,
};
use x_media::media::Media;
#[test]
@@ -559,7 +832,14 @@ mod tests {
"twitter",
"https://x.com/u/status/1",
"My title",
Some(("Author", "https://x.com/u", "My title", "tag1 tag2")),
"My content",
Some((
"Author",
"https://x.com/u",
"My title",
"My content",
"tag1 tag2",
)),
false,
"<a href=\"https://x.com/u\">Author</a> · My title",
&media,
@@ -567,6 +847,7 @@ mod tests {
assert!(report.contains("site: twitter"), "{report}");
assert!(report.contains("key: twitter:1"), "{report}");
assert!(report.contains("title: My title"), "{report}");
assert!(report.contains("content: My content"), "{report}");
assert!(report.contains("author: Author"), "{report}");
assert!(report.contains("author_url: https://x.com/u"), "{report}");
assert!(report.contains("tags: tag1 tag2"), "{report}");
@@ -584,7 +865,7 @@ mod tests {
#[test]
fn debug_report_without_render_data_and_no_media() {
let report = debug_report("u", "pixiv", "s", "t", None, true, "c", &[]);
let report = debug_report("u", "pixiv", "s", "t", "c", None, true, "p", &[]);
assert!(!report.contains("author:"), "{report}");
assert!(report.contains("sensitive: true"), "{report}");
assert!(report.contains("media (0):"), "{report}");
@@ -601,10 +882,12 @@ mod tests {
"twitter",
"https://x.com/u/status/1",
"A & B <C>",
"body & <more>",
Some((
"A &amp; B",
"https://x.com/u",
"A &amp; B &lt;C&gt;",
"body &amp; &lt;more&gt;",
"#a &amp; #b",
)),
false,
@@ -640,8 +923,265 @@ mod tests {
fallback_url: None,
})
.collect();
let report = debug_report("u", "twitter", "s", "t", None, false, "c", &media);
let report = debug_report("u", "twitter", "s", "t", "c", None, false, "p", &media);
assert!(report.chars().count() <= MAX_DEBUG_REPORT_CHARS, "{report}");
assert!(report.ends_with('…'), "{report}");
}
#[test]
fn settings_text_reports_the_chat_configuration() {
use crate::state::ChatData;
// A fresh chat: the defaults must be spelled out, including how to set
// the channel (an empty field is not a status).
let empty = settings_text(&ChatData::default());
assert!(empty.contains("Forward channel: not set"), "{empty}");
assert!(empty.contains("/set_forward_channel"), "{empty}");
assert!(empty.contains("Edit before forward: off"), "{empty}");
assert!(empty.contains("built-in for every site"), "{empty}");
assert!(empty.contains("Templates: none"), "{empty}");
let configured = ChatData {
forward_channel_id: Some(-100123),
edit_before_forward: true,
template: [("b", "[]"), ("a", "[]")]
.into_iter()
.map(|(k, v)| (k.to_string(), v.to_string()))
.collect(),
message_format: [("twitter", "{author}: {content}")]
.into_iter()
.map(|(k, v)| (k.to_string(), v.to_string()))
.collect(),
..ChatData::default()
};
let text = settings_text(&configured);
assert!(text.contains("Forward channel: -100123"), "{text}");
assert!(text.contains("Edit before forward: on"), "{text}");
assert!(text.contains("twitter => {author}: {content}"), "{text}");
// Sorted, so the same chat always reports the same thing.
assert!(text.contains("Templates (2): a, b"), "{text}");
}
#[test]
fn help_and_start_cover_what_the_command_list_cannot() {
// The placeholders the renderer substitutes must be the ones the help
// lists: a stale list is worse than none.
for placeholder in super::FORMAT_PLACEHOLDERS {
assert!(
super::HELP_FOOTER.contains(&format!("{{{placeholder}}}")),
"help does not document {{{placeholder}}}"
);
}
// The private-chat rule and the template placeholder semantics are the
// two things users got wrong most often.
assert!(super::HELP_FOOTER.contains("private chats only"));
assert!(super::HELP_FOOTER.contains("[]"));
assert!(super::START_TEXT.contains("inline mode"));
assert!(super::START_TEXT.contains("/help"));
// Both must stay inside Telegram's message limit.
assert!(super::HELP_FOOTER.chars().count() < 2000);
assert!(super::START_TEXT.chars().count() < 2000);
}
#[test]
fn every_command_is_registered_and_parses() {
use teloxide::utils::command::BotCommands;
use super::Command;
let registered: Vec<String> = Command::bot_commands()
.into_iter()
.map(|command| command.command.trim_start_matches('/').to_string())
.collect();
for expected in [
"start",
"help",
"settings",
"set_forward_channel",
"remove_template",
"set_format",
"test",
"debug",
] {
assert!(
registered.iter().any(|name| name == expected),
"{expected} missing from {registered:?}"
);
}
// Telegram caps a command description at 256 chars.
for command in Command::bot_commands() {
assert!(
command.description.chars().count() <= 256,
"{}: description too long",
command.command
);
}
// A command with a `String` argument must parse with its whole
// argument: without `parse_with`, teloxide's default parser rejects
// `/remove_template x` and the command silently falls through to the
// URL flow.
assert!(matches!(
Command::parse("/settings", ""),
Ok(Command::Settings)
));
match Command::parse("/remove_template tpl", "") {
Ok(Command::RemoveTemplate(name)) => assert_eq!(name, "tpl"),
Ok(_) => panic!("/remove_template parsed as another command"),
Err(e) => panic!("parse error: {e}"),
}
}
#[test]
fn every_documented_invocation_parses() {
use teloxide::utils::command::BotCommands;
use super::Command;
// The README's forms, verbatim. teloxide's `split` parser accepts
// EXACTLY one token per `String` field, so a command documented with
// two arguments (or an optional one) silently stops parsing — and a
// command that does not parse falls through to the URL flow in
// silence.
type Check = fn(&Command) -> bool;
let cases: Vec<(&str, Check)> = vec![
("/start", |c| matches!(c, Command::Start)),
("/help", |c| matches!(c, Command::Help)),
("/settings", |c| matches!(c, Command::Settings)),
("/edit_before_forward", |c| {
matches!(c, Command::EditBeforeForward)
}),
("/remove_forward_channel", |c| {
matches!(c, Command::RemoveForwardChannel)
}),
("/bot_dict", |c| matches!(c, Command::BotDict)),
(
"/set_forward_channel @a_channel",
|c| matches!(c, Command::SetForwardChannel(a) if a == "@a_channel"),
),
(
"/set_template tpl",
|c| matches!(c, Command::SetTemplate(a) if a == "tpl"),
),
(
"/remove_template tpl",
|c| matches!(c, Command::RemoveTemplate(a) if a == "tpl"),
),
(
"/set_format twitter {author}: {title}",
|c| matches!(c, Command::SetFormat(a) if a == "twitter {author}: {title}"),
),
(
"/set_format twitter -",
|c| matches!(c, Command::SetFormat(a) if a == "twitter -"),
),
// Documented as "clear everything" when called without a link.
(
"/clear_cache",
|c| matches!(c, Command::ClearCache(a) if a.is_empty()),
),
(
"/clear_cache https://x.com/u/status/1",
|c| matches!(c, Command::ClearCache(a) if a == "https://x.com/u/status/1"),
),
(
"/test https://x.com/u/status/1",
|c| matches!(c, Command::Test(a) if a == "https://x.com/u/status/1"),
),
(
"/debug https://x.com/u/status/1",
|c| matches!(c, Command::Debug(a) if a == "https://x.com/u/status/1"),
),
];
for (text, ok) in cases {
match Command::parse(text, "") {
Ok(parsed) => assert!(ok(&parsed), "{text} parsed as the wrong variant"),
Err(e) => panic!("{text} did not parse: {e}"),
}
}
}
#[test]
fn preview_caption_applies_the_chat_format_and_the_long_post_quote() {
let fields = Some((
"Author",
"https://x.com/u",
"Pinned title",
"Pinned body",
"#tag",
));
// No format override → the site's built-in caption, untouched.
assert_eq!(
preview_caption(
"",
"built-in caption",
"https://x.com/u/status/1",
fields,
200
),
"built-in caption"
);
// The bug this pins: `/debug` used to print the built-in caption even
// with a format set, so `/set_format` looked like it did nothing.
let formatted = preview_caption(
"{author} · {title}",
"built-in caption",
"https://x.com/u/status/1",
fields,
200,
);
assert_eq!(formatted, "Author · Pinned title");
// `{url}` comes from the canonical post URL, as in the send paths.
assert_eq!(
preview_caption(
"{url} {title}",
"built-in",
"https://x.com/u/status/1",
fields,
200
),
"https://x.com/u/status/1 Pinned title"
);
// A long post's text is quoted exactly like the send paths quote it.
let long = "".repeat(300);
let fields = Some(("Author", "https://x.com/u", "", long.as_str(), ""));
let quoted = preview_caption(
"",
"https://x.com/u/status/1\n<a href=\"https://x.com/u\">Author</a>: 正…",
"https://x.com/u/status/1",
fields,
200,
);
assert!(quoted.contains("<blockquote expandable>"), "{quoted}");
// Without render fields (a site that does not expose them) the
// built-in caption is all there is.
assert_eq!(
preview_caption("", "built-in", "https://x.com/u/status/1", None, 200),
"built-in"
);
}
#[test]
fn unknown_placeholder_finds_typos_only() {
assert_eq!(unknown_placeholder("{author} — {title}"), None);
// Every key the renderer substitutes must pass, in any combination.
assert_eq!(
unknown_placeholder("{url}{author}{author_url}{title}{content}{tags}"),
None
);
// Plain text and braces Telegram renders literally are not tokens.
assert_eq!(unknown_placeholder("no placeholders here"), None);
assert_eq!(unknown_placeholder("{unclosed"), None);
assert_eq!(unknown_placeholder("{titel}"), Some("titel"));
assert_eq!(unknown_placeholder("{title} {Content}"), Some("Content"));
// A typo after a valid token is still found.
assert_eq!(unknown_placeholder("{url} {tag}"), Some("tag"));
}
}
+26 -2
View File
@@ -127,10 +127,34 @@ async fn answer_inline_query(bot: Bot, query: InlineQuery) -> Result<bool, Reque
Ok(Some(fetched)) => {
let mut results: Vec<InlineQueryResult> = Vec::new();
// Inline results have the same 1024-char caption limit as regular
// messages; truncate once here for all items.
// messages; truncate once here for all items, then apply the same
// long-post quoting as the send paths. `answer_inline_query` has no
// `AppContext` (the debounce spawns it), so the parsed config comes
// from the process-wide static, and the text is the *escaped*
// title/content the built-in caption embeds (the raw
// `Fetched.title`/`content` differ whenever the post contains
// `<`/`&`).
let caption = x_media::site::truncate_caption(&fetched.caption);
let text = fetched
.render_fields()
.map(|(_, _, title, content, _)| x_media::site::compose_text(title, content))
.unwrap_or_default();
let caption = crate::send::quote_long_caption(
&caption,
&text,
super::CONFIG.caption_quote_text_chars,
);
for (i, media) in fetched.media.iter().enumerate() {
let id = format!("{i}");
// Telegram fetches an inline result's URL itself and cannot
// send site-specific headers, so hotlink-protected media
// (pixiv's pximg.net) would render as a broken file there.
// Locally produced media (ugoira MP4, bsky remux) is a local
// path and does not parse as a URL at all — same skip.
if x_media::site::needs_media_headers(media.url()) {
log::debug!("inline: skipping hotlink-protected media {id}");
continue;
}
let Some(url) = url::Url::parse(media.url()).ok() else {
continue;
};
@@ -138,7 +162,7 @@ async fn answer_inline_query(bot: Bot, query: InlineQuery) -> Result<bool, Reque
.thumbnail_url()
.and_then(|t| url::Url::parse(t).ok())
.unwrap_or_else(|| url.clone());
let caption = caption.clone();
let caption = caption.clone().into_owned();
let result = match media {
Media::Illustration { .. } => {
// Inline photo results have their own (smaller) size
+66 -3
View File
@@ -23,7 +23,9 @@ use crate::media_sender::MediaSender;
use commands::{Command, execute_command};
use teloxide::RequestError;
use teloxide::prelude::*;
use teloxide::types::{ChatId, ChatKind, Message, MessageId, ParseMode, ReplyParameters};
use teloxide::types::{
ChatId, ChatKind, Message, MessageId, ParseMode, PublicChatKind, ReplyParameters,
};
use teloxide::utils::command::BotCommands;
use urls::{URL_JOBS, extract_urls};
@@ -60,8 +62,8 @@ pub(crate) async fn reply_html(
/// Log prefix tying the whole lifecycle of one link (fetch → send → cache →
/// forward) together: the normalized cache key (`twitter:123…`, `pixiv:123`,
/// `bsky:handle/rkey`) instead of the raw URL, so logs stay short and do not
/// echo full user-submitted URLs at info level.
/// `bsky:handle/rkey`, `bilibili:123…`) instead of the raw URL, so logs stay
/// short and do not echo full user-submitted URLs at info level.
pub fn log_key(url: &str) -> String {
x_media::site::cache_key(url).unwrap_or_else(|| "<unsupported>".to_string())
}
@@ -175,10 +177,38 @@ pub async fn message_handler(bot: Bot, message: Message) -> Result<(), RequestEr
break;
}
}
} else if is_group(&message.chat.kind)
&& extract_urls(&message)
.iter()
.any(|url| x_media::site::cache_key(url).is_some())
{
// A supported link in a group used to be dropped in silence, which
// reads as a broken bot (the command menu is registered globally, so
// the expectation is there). Unsupported links stay ignored; the hint
// names the two paths that do work. Channels are excluded — the reply
// would be posted into the channel itself.
let _ = reply(&bot, message.chat.id.0, message.id, GROUP_LINK_HINT).await;
}
respond(())
}
/// Answer for a link posted where the pipeline does not run (a group): links
/// are private-chat only, inline mode is the group path.
const GROUP_LINK_HINT: &str =
"Links are handled in private chat only — send me this link there, or use inline mode here.";
/// Groups and supergroups, as opposed to private chats and channels.
fn is_group(kind: &ChatKind) -> bool {
matches!(
kind,
ChatKind::Public(chat)
if matches!(
chat.kind,
PublicChatKind::Group | PublicChatKind::Supergroup(_)
)
)
}
#[cfg(test)]
mod tests {
use super::*;
@@ -271,4 +301,37 @@ mod tests {
assert!(!edit_message_handler(&ctx, 1, PROMPT_ID, "hello").await);
assert!(sender.calls().is_empty());
}
#[test]
fn the_link_hint_is_for_groups_only() {
use teloxide::types::{ChatPrivate, ChatPublic, PublicChatChannel, PublicChatSupergroup};
let group = ChatKind::Public(ChatPublic {
title: None,
kind: PublicChatKind::Group,
});
let supergroup = ChatKind::Public(ChatPublic {
title: None,
kind: PublicChatKind::Supergroup(PublicChatSupergroup {
username: None,
is_forum: false,
}),
});
// A channel must stay silent: the hint reply would be posted into the
// channel itself.
let channel = ChatKind::Public(ChatPublic {
title: None,
kind: PublicChatKind::Channel(PublicChatChannel { username: None }),
});
let private = ChatKind::Private(ChatPrivate {
username: None,
first_name: None,
last_name: None,
});
assert!(is_group(&group));
assert!(is_group(&supergroup));
assert!(!is_group(&channel));
assert!(!is_group(&private));
}
}
+278 -26
View File
@@ -4,9 +4,11 @@
use super::{log_key, reply};
use crate::ctx::{AppContext, CONTEXT};
use crate::link_cache::{CachedMediaKind, CachedPost};
use crate::media_sender::MediaSender;
use crate::send::{self, MediaItemPayload, Task};
use crate::state::ChatData;
use std::collections::HashSet;
use std::future::Future;
use std::sync::LazyLock;
use teloxide::types::{ChatAction, ChatId, Message, MessageEntityKind, MessageId};
use x_media::media::Media;
@@ -190,11 +192,16 @@ async fn dispatch_send(
log_key(url)
);
send::enqueue_retry(ctx.task_queue, *task, delay_seconds).await;
// Name the post and the wait: "queued for retry" alone left the
// user guessing which link it was and how long the wait is.
let _ = reply(
ctx.sender,
chat_id,
reply_to,
"Send failed. Task queued for retry.",
format!(
"Send failed for {} — retrying in {delay_seconds:.0}s.",
log_key(url)
),
)
.await;
}
@@ -285,6 +292,11 @@ fn build_send_task(
/// workers pass [`PostSend::FromChat`], the `/test` command
/// [`PostSend::Suppressed`]. Everything else (cache write, retry enqueue,
/// dead-letter notification) is identical.
///
/// Wraps [`url_media_inner`] with the chat-action keep-alive: Telegram expires
/// an action indicator after ~5s, while a fetch (ugoira encode, HLS remux) plus
/// a download-and-reupload fallback routinely takes longer — without the
/// refresh the chat shows nothing and the bot reads as stalled.
pub(crate) async fn url_media(
ctx: &AppContext<'_>,
chat_id: i64,
@@ -292,14 +304,136 @@ pub(crate) async fn url_media(
url: &str,
post_send: PostSend,
) {
let reply_to = MessageId(reply_to_message_id as i32);
if let Err(e) = ctx
.sender
.send_chat_action(ChatId(chat_id), ChatAction::Typing)
.await
{
// Shared with the pipeline: once the media types are known the indicator
// switches from "typing" to "sending photo/video".
let hint = parking_lot::Mutex::new(ActionHint::Typing);
run_with_chat_action(
ctx.sender,
chat_id,
&hint,
url_media_inner(ctx, chat_id, reply_to_message_id, url, post_send, &hint),
)
.await;
}
/// Runs `pipeline` while keeping the chat's action indicator alive: Telegram
/// expires an action after ~5s, while a fetch (ugoira encode, HLS remux) plus a
/// download-and-reupload fallback routinely takes longer. The pipeline updates
/// `hint` when it knows what it is sending.
async fn run_with_chat_action<F: Future<Output = ()>>(
sender: &dyn MediaSender,
chat_id: i64,
hint: &parking_lot::Mutex<ActionHint>,
pipeline: F,
) {
// The guard is released before the await: a parking_lot guard held across
// it makes the future !Send, and the URL workers spawn these.
let action = hint.lock().action();
if let Err(e) = sender.send_chat_action(ChatId(chat_id), action).await {
log::error!("send_chat_action failed: {e}");
}
tokio::pin!(pipeline);
loop {
tokio::select! {
// `biased` polls the pipeline first, so a finished pipeline returns
// without ever arming the refresh timer (no stray actions).
biased;
() = &mut pipeline => return,
() = tokio::time::sleep(ACTION_REFRESH) => {
let action = hint.lock().action();
if let Err(e) = sender.send_chat_action(ChatId(chat_id), action).await {
log::error!("send_chat_action failed: {e}");
}
}
}
}
}
/// How often the chat-action indicator is refreshed while a pipeline runs.
/// Telegram's indicator lasts ~5s; refreshing slightly inside that keeps it
/// on-screen continuously.
const ACTION_REFRESH: std::time::Duration = std::time::Duration::from_secs(4);
/// What the chat action should say. Unknown before the fetch, so the pipeline
/// starts with `Typing` and switches as soon as the media types are known.
#[derive(Clone, Copy)]
enum ActionHint {
Typing,
Photo,
Video,
}
impl ActionHint {
/// Photos make Telegram label the send "sending photo"; video/animation
/// only payloads get "sending video". A mixed post takes the photo label
/// (the group's first item is always a photo, see `photos_first`).
fn for_items(items: &[MediaItemPayload]) -> Self {
if items
.iter()
.any(|item| matches!(item, MediaItemPayload::Photo { .. }))
{
Self::Photo
} else {
Self::Video
}
}
fn action(self) -> ChatAction {
match self {
Self::Typing => ChatAction::Typing,
Self::Photo => ChatAction::UploadPhoto,
Self::Video => ChatAction::UploadVideo,
}
}
}
/// User-facing text for a failed fetch. The [`FetchError`] class is what tells
/// the user whether the post is gone, withheld or the source is refusing
/// requests; a single generic sentence threw that away.
fn fetch_error_message(err: &x_media::site::FetchError) -> String {
use x_media::site::FetchError;
match err {
FetchError::NotFound => "Post not found (deleted, private or unavailable).".to_string(),
FetchError::Sensitive => concat!(
"This post's media is withheld (age-restricted). ",
"The bot owner must set TWITTER_AUTH_TOKEN to fetch it."
)
.to_string(),
FetchError::Blocked => {
"The source site refused the request (risk control). Try again later.".to_string()
}
FetchError::Disabled { site } => {
format!("{} support is disabled on this bot.", site_title(site))
}
FetchError::Transient(_) | FetchError::Http(_) => {
"The source site is unavailable right now (tried 3 times). Try again later.".to_string()
}
// Parse/shape surprises, pixiv auth details, oversized media: nothing
// actionable for the user beyond "this did not work".
_ => "Failed to fetch media from this link.".to_string(),
}
}
/// Site ids are lowercase ASCII (`pixiv`); user-facing text capitalizes the
/// first letter.
fn site_title(site: &str) -> String {
let mut chars = site.chars();
match chars.next() {
Some(first) => first.to_uppercase().collect::<String>() + chars.as_str(),
None => String::new(),
}
}
#[allow(clippy::too_many_arguments)]
async fn url_media_inner(
ctx: &AppContext<'_>,
chat_id: i64,
reply_to_message_id: i64,
url: &str,
post_send: PostSend,
hint: &parking_lot::Mutex<ActionHint>,
) {
let reply_to = MessageId(reply_to_message_id as i32);
// Link cache: a post sent before is re-sent from Telegram file ids —
// no source-site request, no download, no upload. Keyed by the
@@ -327,6 +461,7 @@ pub(crate) async fn url_media(
&cached.author,
&cached.author_url,
&cached.title,
&cached.content,
&cached.tags,
)
};
@@ -354,6 +489,9 @@ pub(crate) async fn url_media(
},
})
.collect();
// The indicator switches to "sending photo/video" once the kinds are
// known; `items` is moved into the task below.
*hint.lock() = ActionHint::for_items(&items);
let task = build_send_task(
&chat_data,
chat_id,
@@ -377,13 +515,7 @@ pub(crate) async fn url_media(
// Retries exhausted: notify the user (Rust-only requirement 3).
Err(e) => {
log::error!("fetch {url}: {e}");
let _ = reply(
ctx.sender,
chat_id,
reply_to,
"Failed to fetch media from this link.",
)
.await;
let _ = reply(ctx.sender, chat_id, reply_to, fetch_error_message(&e)).await;
}
Ok(Some(mut fetched)) => {
if fetched.media.is_empty() {
@@ -406,23 +538,28 @@ pub(crate) async fn url_media(
let caption = fetched.caption_with(&format);
// Raw render data for the link cache; the send fills in the
// Telegram file ids and persists the entry.
let cache_data = fetched
.render_fields()
.map(|(author, author_url, title, tags)| CachedPost {
url: fetched.source_url.clone(),
caption: fetched.caption.clone(),
title: title.to_string(),
author: author.to_string(),
author_url: author_url.to_string(),
tags: tags.to_string(),
sensitive: fetched.sensitive,
media: vec![],
});
let cache_data =
fetched
.render_fields()
.map(|(author, author_url, title, content, tags)| CachedPost {
url: fetched.source_url.clone(),
caption: fetched.caption.clone(),
title: title.to_string(),
content: content.to_string(),
author: author.to_string(),
author_url: author_url.to_string(),
tags: tags.to_string(),
sensitive: fetched.sensitive,
media: vec![],
});
let items: Vec<MediaItemPayload> = fetched
.media
.iter()
.map(|media| media_to_payload(media, fetched.sensitive))
.collect();
// The indicator switches to "sending photo/video" once the kinds
// are known; `items` is moved into the task below.
*hint.lock() = ActionHint::for_items(&items);
let task = build_send_task(
&chat_data,
chat_id,
@@ -465,6 +602,7 @@ mod tests {
url: "https://x.com/u/status/1".into(),
caption: "cap".into(),
title: "t".into(),
content: "c".into(),
author: "a".into(),
author_url: "au".into(),
tags: "".into(),
@@ -530,6 +668,35 @@ mod tests {
);
}
/// The caption-quote threshold matches the post's text inside the caption,
/// so a long-text cache hit is quoted and a short-text one is not.
#[tokio::test]
async fn cache_hit_quotes_a_long_text_caption() {
let mut stores = TestStores::new();
stores.config_mut().caption_quote_text_chars = 3;
let prefix = "https://x.com/u/status/1\n<a href=\"au\">a</a>: ";
for (text, expected) in [
(
"abc",
format!("{prefix}<blockquote expandable>abc</blockquote>"),
),
("ab", format!("{prefix}ab")),
] {
let sender = MockSender::scripted(vec![Outcome::GroupOk], permanent_error);
let ctx = stores.ctx(&sender);
let mut entry = cached_photo_entry();
entry.caption = format!("{prefix}{text}");
entry.title = String::new();
entry.content = text.into();
stores.link_cache().put("twitter:1", &entry).await;
url_media(&ctx, 1, 2, "https://x.com/u/status/1", PostSend::FromChat).await;
assert_eq!(sender.captions(), vec![expected], "text {text:?}");
}
}
#[tokio::test]
async fn unsupported_url_is_ignored_silently() {
let stores = TestStores::new();
@@ -662,4 +829,89 @@ mod tests {
// Dead-letter notification still reaches the chat that asked.
assert_eq!(notify_chat_id, Some(1));
}
#[test]
fn fetch_errors_map_to_distinct_user_messages() {
use x_media::site::FetchError;
let disabled = fetch_error_message(&FetchError::Disabled { site: "pixiv" });
assert_eq!(disabled, "Pixiv support is disabled on this bot.");
assert_eq!(
fetch_error_message(&FetchError::NotFound),
"Post not found (deleted, private or unavailable)."
);
let sensitive = fetch_error_message(&FetchError::Sensitive);
assert!(sensitive.contains("TWITTER_AUTH_TOKEN"), "{sensitive}");
let blocked = fetch_error_message(&FetchError::Blocked);
assert!(blocked.contains("refused"), "{blocked}");
// Each class that has something to say must differ from the generic
// fallback — one generic sentence for everything is what this fixes.
let generic = fetch_error_message(&FetchError::TooLarge);
for text in [disabled, sensitive, blocked] {
assert_ne!(text, generic);
}
}
#[test]
fn action_hint_follows_the_media_kind() {
use MediaItemPayload::{Animation, Photo, Video};
let photo = || Photo {
media: "https://p/1.jpg".into(),
has_spoiler: false,
fallback_url: None,
file_id: false,
};
let video = || Video {
media: "https://v/1.mp4".into(),
has_spoiler: false,
thumbnail: None,
fallback_url: None,
file_id: false,
};
// Unknown before the fetch: the pipeline starts on "typing".
assert!(matches!(ActionHint::Typing.action(), ChatAction::Typing));
assert!(matches!(
ActionHint::for_items(&[photo()]).action(),
ChatAction::UploadPhoto
));
assert!(matches!(
ActionHint::for_items(&[
video(),
Animation {
media: "https://v/2.mp4".into(),
has_spoiler: false,
file_id: false,
}
])
.action(),
ChatAction::UploadVideo
));
// A mixed post takes the photo label: `photos_first` always leads with
// a photo, which is what Telegram shows.
assert!(matches!(
ActionHint::for_items(&[video(), photo()]).action(),
ChatAction::UploadPhoto
));
}
#[tokio::test(start_paused = true)]
async fn a_long_pipeline_keeps_the_chat_action_alive() {
let sender = MockSender::scripted(vec![], permanent_error);
let hint = parking_lot::Mutex::new(ActionHint::Typing);
// Three refresh windows of work: Telegram would have dropped the
// indicator twice without the keep-alive.
let pipeline = async { tokio::time::sleep(ACTION_REFRESH * 3).await };
run_with_chat_action(&sender, 1, &hint, pipeline).await;
let actions = sender
.calls()
.iter()
.filter(|call| **call == "send_chat_action")
.count();
assert_eq!(actions, 3, "expected the initial action plus two refreshes");
}
}
+48
View File
@@ -38,6 +38,10 @@ pub struct CachedPost {
/// override).
pub caption: String,
pub title: String,
/// The post's body text. Defaulted on read: entries written before the
/// title/content split carry it inside `title`.
#[serde(default)]
pub content: String,
pub author: String,
pub author_url: String,
pub tags: String,
@@ -182,6 +186,7 @@ mod tests {
url: "https://x.com/u/status/1".into(),
caption: "cap".into(),
title: "t".into(),
content: "c".into(),
author: "a".into(),
author_url: "au".into(),
tags: "".into(),
@@ -193,6 +198,49 @@ mod tests {
}
}
/// A payload written before the title/content split has no `content`
/// field. It must still read back — the cache deletes what it cannot
/// parse — with its text left where it was stored (`title`) and the
/// caption it replays untouched. No migration: a self-hosted cache entry
/// lives one TTL, and moving the text would only reshuffle `/set_format`
/// placeholders until it expires.
#[tokio::test]
async fn pre_split_entry_still_parses() {
let dir = tempfile::tempdir().unwrap();
let cache = LinkCache::new(
crate::db::open_store(dir.path().join("c.db").to_str().unwrap()).unwrap(),
);
let legacy = serde_json::json!({
"url": "https://x.com/u/status/1",
"caption": "https://x.com/u/status/1\n<a href=\"au\">a</a>: old text",
"title": "old text",
"author": "a",
"author_url": "au",
"tags": "",
"sensitive": false,
"media": [{"kind": "photo", "file_id": "AgAC..."}]
});
{
let conn = rusqlite::Connection::open(dir.path().join("c.db")).unwrap();
conn.execute(
"INSERT INTO link_cache (url, payload, created_at) VALUES (?1, ?2, ?3)",
params!["twitter:1", legacy.to_string(), now_f64()],
)
.unwrap();
}
let got = cache
.get("twitter:1", Duration::from_secs(3600))
.await
.expect("a pre-split payload must not be dropped");
assert_eq!(got.title, "old text");
assert_eq!(got.content, "");
assert_eq!(
got.caption,
"https://x.com/u/status/1\n<a href=\"au\">a</a>: old text"
);
}
#[tokio::test]
async fn put_get_roundtrip() {
let dir = tempfile::tempdir().unwrap();
+10 -4
View File
@@ -2,7 +2,7 @@ use dotenv::dotenv;
use teloxide::dptree::endpoint;
use teloxide::prelude::*;
use teloxide::stop::StopToken;
use teloxide::types::{ChatId, InputFile, MessageId};
use teloxide::types::{ChatId, InlineKeyboardMarkup, InputFile, MessageId};
use teloxide::update_listeners::{self, UpdateListener, webhooks};
use tokio::sync::watch;
use x_media::site;
@@ -119,13 +119,19 @@ async fn main() {
log::debug!("rate limiter: dropped {idle_limiters} idle bucket(s)");
}
for (chat_id, prompt_message_id) in removed {
// If the prompt was already deleted, this fails with a
// 400 "message to edit not found" — log and ignore.
// Rewritten in place, not announced: the sweep is a
// background timer, and a fresh message would wake the chat
// up to a full TTL later about a prompt the user already
// walked away from. The edit drops the buttons too. If the
// prompt was already deleted this fails with a 400
// "message to edit not found" — log and ignore.
if let Err(e) = bot
.edit_message_reply_markup(
.edit_message_text(
ChatId(chat_id),
MessageId(prompt_message_id as i32),
send::EDIT_PROMPT_EXPIRED_TEXT,
)
.reply_markup(InlineKeyboardMarkup::default())
.await
{
log::info!("edit-expiry sweep: prompt message gone: {e}");
+12 -1
View File
@@ -327,9 +327,20 @@ pub(crate) mod test_support {
&self,
_chat_id: ChatId,
_reply_to: MessageId,
_items: Vec<InputMedia>,
items: Vec<InputMedia>,
) -> BoxFuture<'_, Result<Vec<Message>, RequestError>> {
Box::pin(async move {
// Record the captions exactly as Telegram receives them (only
// the first item of a group carries one), so tests can assert
// what a recipient sees.
self.captions
.lock()
.extend(items.iter().filter_map(|item| match item {
InputMedia::Photo(photo) => photo.caption.clone(),
InputMedia::Video(video) => video.caption.clone(),
InputMedia::Animation(animation) => animation.caption.clone(),
_ => None,
}));
match self.next("send_media_group") {
Outcome::GroupOk => Ok(Vec::new()),
Outcome::GroupErr => Err(self.error()),
+10 -2
View File
@@ -342,8 +342,16 @@ impl QueueWorker {
payload,
}) => {
if row.attempts as u32 >= MAX_RETRIES {
let message = format!("task failed after {MAX_RETRIES} retries");
log::error!("dead-lettering {}: {message}", row.id);
// The queue keeps only the payload, not the last error, so
// the cause of an exhausted retry is just that: exhausted.
// (The dead-letter message is read by the user, so it must
// not restate its own wrapper — see `failure_text`.)
let message = "retries exhausted".to_string();
log::error!(
"dead-lettering {}: {message} after {} attempt(s)",
row.id,
row.attempts + 1
);
self.delete_row(&row.id).await;
(self.dead_letter)(payload, message).await;
} else {
+265 -15
View File
@@ -17,8 +17,8 @@ use crate::link_cache::{CachedMedia, CachedMediaKind, CachedPost};
use crate::media_sender::MediaSender;
use input_media::{build_media_group, input_file_for, item_url};
use post_send::{cache_animation_send, cache_sent_task};
use rand::Rng;
use serde::{Deserialize, Serialize};
use std::borrow::Cow;
use std::sync::LazyLock;
use teloxide::prelude::*;
use teloxide::types::{ChatId, InputFile, InputMedia, MessageId};
@@ -28,8 +28,8 @@ use upload::{FallbackError, PreparedItem, prepare_upload_item, send_batch_via_up
// The crate-facing API of this module lives in its submodules; re-export the
// parts other modules use so call sites stay `send::x`.
pub(crate) use post_send::{
KEEP_ALIVE, Settled, dead_letter_notify, enqueue_retry, handle_task, post_send_actions,
settle_task,
EDIT_PROMPT_EXPIRED_TEXT, KEEP_ALIVE, Settled, dead_letter_notify, enqueue_retry, handle_task,
post_send_actions, settle_task,
};
/// One process-wide Bot for queue workers. Building a fresh Bot (and its HTTP
@@ -230,7 +230,9 @@ fn collect_file_ids(messages: &[Message], batch: &[MediaItemPayload], out: &mut
}
}
pub const MAX_MEDIA_GROUP: usize = 9;
/// Telegram's `sendMediaGroup` accepts 210 items per group; 10 (not the older
/// 9) means a 10-image post arrives as one album instead of two messages.
pub const MAX_MEDIA_GROUP: usize = 10;
/// Splits media into batches of at most [`MAX_MEDIA_GROUP`] items, moving the
/// items out (no per-item clone).
@@ -261,7 +263,7 @@ pub fn photos_first(items: Vec<MediaItemPayload>) -> Vec<MediaItemPayload> {
/// Exponential backoff with jitter, capped at 30s.
pub fn retry_delay_seconds(attempts: u32) -> f64 {
let jitter: f64 = rand::thread_rng().gen_range(0.2..0.8);
let jitter: f64 = rand::random_range(0.2..0.8);
(2f64.powi(attempts as i32) + jitter).min(30.0)
}
@@ -389,6 +391,61 @@ fn updated_sequence_task(task: &Task, batch_index: usize, sent_message_ids: Vec<
updated
}
/// The caption's text tail: everything after the author link, provided it
/// really is the post's text.
///
/// `text` is the *escaped* title + content the caption embeds; the caption may
/// have been truncated inside it, in which case only its prefix is present, so
/// the tail only has to match the text's start. `None` for a caption with
/// another layout — pixiv's title-inside-a-link, a `/set_format` that moves
/// `{title}`/`{content}` off the author line — which is left unquoted instead
/// of guessing where the text begins.
fn text_tail<'c>(caption: &'c str, text: &str) -> Option<&'c str> {
let (_, tail) = caption.rsplit_once("</a>: ")?;
let visible = tail.strip_suffix('\u{2026}').unwrap_or(tail);
(!visible.is_empty() && text.starts_with(visible)).then_some(tail)
}
/// The text a task's caption embeds, read from the same cache snapshot the
/// caption came from: `title` and `content` joined the way the sites' built-in
/// captions join them.
fn task_text(task: &Task) -> String {
task.cache_data()
.map(|data| x_media::site::compose_text(&data.title, &data.content))
.unwrap_or_default()
}
/// Wraps the post's text inside the caption in an expandable blockquote once
/// that text is long enough that the message would otherwise be a wall of text
/// (`threshold` is `CAPTION_QUOTE_TEXT_CHARS`; `0` disables the wrap). The URL
/// and the author line stay outside the quote.
///
/// Applied at the send boundary, after the caller's `truncate_caption`:
/// Telegram measures a caption *after entities parsing*, so the tags cost no
/// length and a wrapped caption cannot exceed the 1024-character limit.
/// Retries replay the task's (unwrapped) caption, so the decision is remade on
/// every attempt — changing the threshold takes effect immediately.
///
/// A caption that already carries a blockquote is left as it is: the API
/// rejects nested ones ("all other entities can't contain each other"), and a
/// user-written `/set_format` template may contain one.
pub(crate) fn quote_long_caption<'a>(
caption: &'a str,
text: &str,
threshold: usize,
) -> Cow<'a, str> {
if threshold == 0 || caption.contains("<blockquote") || text.chars().count() < threshold {
return Cow::Borrowed(caption);
}
let Some(tail) = text_tail(caption, text) else {
return Cow::Borrowed(caption);
};
let prefix = &caption[..caption.len() - tail.len()];
Cow::Owned(format!(
"{prefix}<blockquote expandable>{tail}</blockquote>"
))
}
/// Sends the media batches starting at `task.batch_index`, extending
/// `sent_message_ids`. Returns all sent message ids on full success; on
/// failure returns a [`SendError`] whose task carries the resumed state.
@@ -407,6 +464,10 @@ pub async fn send_media_sequence(ctx: &AppContext<'_>, task: &Task) -> Result<Ve
};
let chat_id = *chat_id;
let reply_to = *reply_to_message_id;
// A long post is quoted so the message reads as a card rather than a wall
// of text; the text comes from the same cache snapshot as the caption.
let text = task_text(task);
let caption = quote_long_caption(caption, &text, ctx.config.caption_quote_text_chars);
let mut sent = sent_message_ids.clone();
// File ids accumulated across batches for the link cache. Only a fresh
// (non-resumed) full send populates the cache.
@@ -415,7 +476,7 @@ pub async fn send_media_sequence(ctx: &AppContext<'_>, task: &Task) -> Result<Ve
for idx in *batch_index..media_batches.len() {
let batch = &media_batches[idx];
let caption = if idx == 0 {
Some(caption.as_str())
Some(caption.as_ref())
} else {
None
};
@@ -516,6 +577,10 @@ pub async fn send_animation(ctx: &AppContext<'_>, task: &Task) -> Result<Vec<i64
};
let chat_id = *chat_id;
let reply_to = *reply_to_message_id;
// Same long-post quoting as the media-group path (see
// `quote_long_caption`); both sends below share this string.
let text = task_text(task);
let caption = quote_long_caption(caption, &text, ctx.config.caption_quote_text_chars);
let (media_url, has_spoiler) = match animation {
MediaItemPayload::Animation {
media, has_spoiler, ..
@@ -537,7 +602,7 @@ pub async fn send_animation(ctx: &AppContext<'_>, task: &Task) -> Result<Vec<i64
ctx.sender,
chat_id,
reply_to,
caption,
&caption,
has_spoiler,
url_file,
)
@@ -572,7 +637,7 @@ pub async fn send_animation(ctx: &AppContext<'_>, task: &Task) -> Result<Vec<i64
ctx.sender,
chat_id,
reply_to,
caption,
&caption,
has_spoiler,
animation.media,
)
@@ -665,13 +730,14 @@ mod tests {
fn chunk_media_items_sizes() {
assert_eq!(chunk_media_items::<i32>(vec![]), Vec::<Vec<i32>>::new());
assert_eq!(chunk_media_items((0..9).collect()).len(), 1);
assert_eq!(chunk_media_items((0..10).collect()).len(), 2);
assert_eq!(chunk_media_items((0..10).collect()).len(), 1);
assert_eq!(chunk_media_items((0..11).collect()).len(), 2);
assert_eq!(chunk_media_items((0..25).collect()).len(), 3);
assert_eq!(chunk_media_items((0..25).collect())[2].len(), 7);
assert_eq!(chunk_media_items((0..25).collect())[2].len(), 5);
assert!(
chunk_media_items((0..25).collect())
.iter()
.all(|c| c.len() <= 9)
.all(|c| c.len() <= MAX_MEDIA_GROUP)
);
}
@@ -744,7 +810,88 @@ mod tests {
.flatten()
.map(|button| button.text.clone())
.collect();
assert_eq!(labels, ["a", "b", "m", "q", "y", "z", "↩️ Confirm"]);
assert_eq!(
labels,
["a", "b", "m", "q", "y", "z", "↩️ Confirm", "🛑 Skip"]
);
}
#[test]
fn edit_markup_folds_and_caps_the_template_buttons() {
// Telegram rejects a keyboard over 100 buttons, which would drop the
// whole prompt; the cap keeps it well under that.
let templates: HashMap<String, String> = (0..200)
.map(|i| (format!("t{i:03}"), "[]".to_string()))
.collect();
let keyboard = build_edit_markup(&templates);
let buttons: usize = keyboard.inline_keyboard.iter().map(Vec::len).sum();
assert!(
buttons <= 100,
"a keyboard Telegram rejects would lose the prompt: {buttons}"
);
assert_eq!(
buttons,
super::post_send::MAX_TEMPLATE_BUTTONS + 2,
"the cap plus the confirm/skip pair"
);
// Names are folded, not one per row.
assert_eq!(keyboard.inline_keyboard[0].len(), 3);
assert_eq!(keyboard.inline_keyboard.last().unwrap().len(), 2);
assert_eq!(super::post_send::hidden_template_count(&templates), 140);
// Under the cap nothing is hidden and every name gets a button.
let few: HashMap<String, String> = (0..4)
.map(|i| (format!("t{i}"), "[]".to_string()))
.collect();
assert_eq!(super::post_send::hidden_template_count(&few), 0);
assert_eq!(
build_edit_markup(&few)
.inline_keyboard
.iter()
.map(Vec::len)
.sum::<usize>(),
6
);
}
#[test]
fn failure_text_names_the_post_and_the_cause() {
// A send failure names the post (the cache key) and the cause, so the
// user knows which of their links died.
let task = sequence_task("https://x.com/u/status/1");
let text = super::post_send::failure_text(Some(&task), "retries exhausted");
assert!(text.contains("twitter:1"), "{text}");
assert!(text.contains("retries exhausted"), "{text}");
// A channel-forward failure has no source URL: it must not claim a
// post failed.
let forward = Task::ForwardMessages {
from_chat_id: 1,
to_chat_id: 2,
message_ids: vec![1],
notify_chat_id: None,
notify_message_id: None,
};
let text = super::post_send::failure_text(Some(&forward), "chat not found");
assert!(text.starts_with("Forward failed permanently"), "{text}");
assert!(text.contains("chat not found"), "{text}");
}
#[test]
fn edit_prompt_text_states_the_ttl_and_the_confirm_requirement() {
use std::time::Duration;
let text = super::post_send::edit_prompt_text(Duration::from_secs(24 * 3600));
assert!(text.contains("Expires in 24h"), "{text}");
assert!(text.contains("Confirm"), "{text}");
// The wording of the whole point: no Confirm, no forward.
assert!(text.contains("Nothing is forwarded"), "{text}");
// Sub-hour TTLs must not render "0h".
assert!(
super::post_send::edit_prompt_text(Duration::from_secs(90)).contains("Expires in 1m")
);
assert!(
super::post_send::edit_prompt_text(Duration::from_secs(30)).contains("Expires in 30s")
);
}
#[test]
@@ -931,10 +1078,16 @@ mod tests {
}
fn sequence_task(media: &str) -> Task {
sequence_task_with(media, "cap", None)
}
/// A media-group task; `text` (when given) rides in the link-cache
/// snapshot as `content`, which is where the quote threshold reads it.
fn sequence_task_with(media: &str, caption: &str, text: Option<&str>) -> Task {
Task::SendMediaSequence {
chat_id: 1,
reply_to_message_id: 2,
caption: "cap".into(),
caption: caption.into(),
media_batches: vec![vec![MediaItemPayload::Photo {
media: media.to_string(),
has_spoiler: false,
@@ -948,7 +1101,99 @@ mod tests {
forward_channel_id: None,
notify_chat_id: Some(1),
notify_message_id: Some(2),
cache_data: None,
// The snapshot splits the post's text into title/content the way a
// real fetch does; the quote threshold joins them again.
cache_data: text.map(|text| CachedPost {
url: "https://x.com/u/status/1".into(),
caption: caption.into(),
title: String::new(),
content: text.into(),
author: "me".into(),
author_url: "https://x.com/u".into(),
tags: String::new(),
sensitive: false,
media: vec![],
}),
}
}
/// A media-group task whose built-in caption carries `text` behind the
/// author link — the shape the quote threshold locates the text in.
fn sequence_task_with_text(media: &str, text: &str) -> Task {
let caption =
format!("https://x.com/u/status/1\n<a href=\"https://x.com/u\">me</a>: {text}");
sequence_task_with(media, &caption, Some(text))
}
#[test]
fn quote_long_caption_wraps_only_the_text_tail() {
let text = "一二三四五";
let prefix = "https://x.com/u/status/1\n<a href=\"https://x.com/u\">me</a>: ";
let caption = format!("{prefix}{text}");
// Only the text goes inside the quote; the URL and author line stay
// outside.
assert_eq!(
quote_long_caption(&caption, text, 5),
format!("{prefix}<blockquote expandable>{text}</blockquote>")
);
// One char below the threshold, disabled, and a short text: untouched.
assert_eq!(
quote_long_caption(&caption, text, 6),
format!("{prefix}{text}")
);
assert_eq!(quote_long_caption(&caption, text, 0), caption);
// No author-line anchor means no text to locate — a pixiv caption
// (title inside the link) and a `{content}`-first format stay as they
// are rather than risking a blockquote nested in a tag.
let pixiv =
format!("<a href=\"https://pixiv.net/1\">{text}</a> / <a href=\"u\">me</a>\ntag");
assert_eq!(quote_long_caption(&pixiv, text, 5), pixiv);
let content_first = format!("{text}\nhttps://x.com/u/status/1");
assert_eq!(quote_long_caption(&content_first, text, 5), content_first);
// An empty body has nothing to quote.
assert_eq!(quote_long_caption(prefix, text, 5), prefix);
// A caption that already carries a blockquote is never nested.
let quoted = format!("<blockquote>{caption}</blockquote>");
assert_eq!(quote_long_caption(&quoted, text, 5), quoted);
}
/// `truncate_caption` cuts inside the text and appends an ellipsis; the
/// visible prefix still marks it, so the long-text case that most needs
/// quoting is still quoted.
#[test]
fn quote_long_caption_wraps_a_truncated_text() {
let text = "一二三四五六七八九十";
let prefix = "https://x.com/u/status/1\n<a href=\"https://x.com/u\">me</a>: ";
let caption = format!("{prefix}一二三四五…");
assert_eq!(
quote_long_caption(&caption, text, 5),
format!("{prefix}<blockquote expandable>一二三四五…</blockquote>")
);
}
#[tokio::test]
async fn long_text_caption_reaches_telegram_quoted() {
// The threshold is pinned here instead of read from the environment.
let dir = tempfile::tempdir().unwrap();
let file = dir.path().join("media.jpg");
std::fs::write(&file, b"not-a-real-jpeg").unwrap();
let mut stores = TestStores::new();
stores.config_mut().caption_quote_text_chars = 5;
let prefix = "https://x.com/u/status/1\n<a href=\"https://x.com/u\">me</a>: ";
for (text, expected) in [
(
"一二三四五",
format!("{prefix}<blockquote expandable>一二三四五</blockquote>"),
),
("一二三四", format!("{prefix}一二三四")),
] {
let sender = MockSender::scripted(vec![Outcome::GroupOk], media_fetch_error);
let ctx = stores.ctx(&sender);
let task = sequence_task_with_text(file.to_str().unwrap(), text);
assert!(send_media_sequence(&ctx, &task).await.is_ok());
assert_eq!(sender.captions(), vec![expected], "text {text:?}");
}
}
@@ -1150,7 +1395,11 @@ mod tests {
post_send_actions(&ctx, &task, vec![10, 11]).await;
assert_eq!(sender.calls(), vec!["send_message"]);
assert_eq!(sender.messages(), vec!["Reply to edit message."]);
// The prompt explains the Confirm requirement and the TTL (see the
// pure `edit_prompt_text` test for the exact wording).
let prompt_text = sender.messages().first().cloned().unwrap_or_default();
assert!(prompt_text.contains("Expires in"), "{prompt_text}");
assert!(prompt_text.contains("Confirm"), "{prompt_text}");
// The prompt's own message id keys the record the reply will edit.
let data = stores.chat_store().get(1).await;
let record = data
@@ -1260,6 +1509,7 @@ mod tests {
url: "https://x.com/u/status/1".into(),
caption: "cap".into(),
title: "t".into(),
content: "c".into(),
author: "a".into(),
author_url: "au".into(),
tags: String::new(),
+90 -19
View File
@@ -103,26 +103,76 @@ pub(crate) fn release_keep_alive(task: &Task) {
});
}
/// One button per template name (column layout), then the confirm button.
/// Sorted by name: the templates live in a `HashMap`, so an unsorted walk
/// would reshuffle the buttons between prompts.
/// The edit-before-forward prompt's text. It names both controls and the TTL,
/// because the buttons alone left users waiting for a forward that never came
/// (nothing is forwarded until Confirm).
pub(super) fn edit_prompt_text(ttl: std::time::Duration) -> String {
format!(
"Reply to edit the caption, or tap a template, then ↩️ Confirm to forward. \
Expires in {}. Nothing is forwarded until you confirm.",
coarsest_unit(ttl)
)
}
/// Text the prompt is rewritten to once its record expires. The sweep edits
/// the prompt in place (see `main`): announcing the expiry with a new message
/// would wake the chat up to a full TTL later about a prompt nobody is
/// waiting on.
pub(crate) const EDIT_PROMPT_EXPIRED_TEXT: &str = "⌛ Expired — nothing was forwarded.";
/// `24h` / `90m` / `45s`: the coarsest whole unit, so the prompt stays short.
fn coarsest_unit(ttl: std::time::Duration) -> String {
let secs = ttl.as_secs();
if secs >= 3600 {
format!("{}h", secs / 3600)
} else if secs >= 60 {
format!("{}m", secs / 60)
} else {
format!("{secs}s")
}
}
/// Templates per keyboard row. Telegram rejects a keyboard with more than 100
/// buttons *outright*, which would silently drop the whole prompt, so the
/// names are folded and capped rather than listed one per row.
pub(super) const TEMPLATE_BUTTONS_PER_ROW: usize = 3;
/// Hard cap on template buttons; the prompt text names the ones not shown.
pub(super) const MAX_TEMPLATE_BUTTONS: usize = 60;
/// Template buttons ([`TEMPLATE_BUTTONS_PER_ROW`] per row, at most
/// [`MAX_TEMPLATE_BUTTONS`]), then the confirm/skip pair. Sorted by name: the
/// templates live in a `HashMap`, so an unsorted walk would reshuffle the
/// buttons between prompts.
pub(super) fn build_edit_markup(templates: &HashMap<String, String>) -> InlineKeyboardMarkup {
let mut names: Vec<&String> = templates.keys().collect();
names.sort();
let mut rows = Vec::with_capacity(names.len() + 1);
for name in names {
rows.push(vec![InlineKeyboardButton::callback(
name.clone(),
format!("template|{name}"),
)]);
let shown = names.len().min(MAX_TEMPLATE_BUTTONS);
let mut rows = Vec::with_capacity(shown / TEMPLATE_BUTTONS_PER_ROW + 2);
for chunk in names[..shown].chunks(TEMPLATE_BUTTONS_PER_ROW) {
rows.push(
chunk
.iter()
.map(|name| {
InlineKeyboardButton::callback(name.as_str(), format!("template|{name}"))
})
.collect(),
);
}
rows.push(vec![InlineKeyboardButton::callback(
"↩️ Confirm",
"forward",
)]);
// Skip exists because the prompt holds the forward hostage until Confirm:
// without it the only escape was deleting the message and waiting out the
// TTL for a forward that then never happens.
rows.push(vec![
InlineKeyboardButton::callback("↩️ Confirm", "forward"),
InlineKeyboardButton::callback("🛑 Skip", "skip"),
]);
InlineKeyboardMarkup::new(rows)
}
/// How many templates the markup could not fit, for the prompt text.
pub(super) fn hidden_template_count(templates: &HashMap<String, String>) -> usize {
templates.len().saturating_sub(MAX_TEMPLATE_BUTTONS)
}
/// Notifies a chat about a dead-lettered task (skips when `notify_chat_id` is
/// absent).
pub(super) async fn notify_failure(
@@ -185,12 +235,21 @@ pub(crate) async fn post_send_actions(ctx: &AppContext<'_>, task: &Task, message
};
if edit_before_forward {
let keyboard = build_edit_markup(&ctx.chat_store.get(chat_id).await.template);
let templates = ctx.chat_store.get(chat_id).await.template;
let keyboard = build_edit_markup(&templates);
let mut text = edit_prompt_text(ctx.config.edit_message_ttl);
let hidden = hidden_template_count(&templates);
if hidden > 0 {
// The keyboard is capped; say so instead of silently hiding them.
text.push_str(&format!(
"\n({hidden} more templates not shown — /remove_template to prune.)"
));
}
let prompt = ctx
.sender
.send_message(
ChatId(chat_id),
"Reply to edit message.".to_string(),
text,
Some(MessageId(reply_to as i32)),
Some(keyboard),
)
@@ -247,7 +306,7 @@ pub(crate) async fn post_send_actions(ctx: &AppContext<'_>, task: &Task, message
ctx.sender,
notify_chat_id,
notify_message_id,
&format!("Task failed after retries: {message}"),
&failure_text(None, &message),
)
.await;
}
@@ -341,6 +400,17 @@ async fn send_media_or_animation(ctx: &AppContext<'_>, task: &Task) -> Result<Ve
}
}
/// User-facing text for a task that will never run again: which link died and
/// why. The raw error alone left the user guessing which post it was about.
pub(super) fn failure_text(task: Option<&Task>, message: &str) -> String {
match task.and_then(|task| task.source_url()).map(log_key) {
Some(key) => format!("Send failed permanently for {key}: {message}"),
// `ForwardMessages` carries no source URL: that failure is about the
// channel copy, not about a post.
None => format!("Forward failed permanently: {message}"),
}
}
/// Dead-letter callback wired to the queue in main: settles the task and
/// notifies its chat.
pub(crate) async fn dead_letter_notify(
@@ -351,8 +421,9 @@ pub(crate) async fn dead_letter_notify(
// A dead-lettered task never runs again, and the queue dead-letters retry
// exhaustion itself (the handler is not called again), so this is the only
// place that sees the final payload.
if let Ok(task) = serde_json::from_value::<Task>(payload.clone()) {
settle_task(ctx, &task, Settled::Failed).await;
let task = serde_json::from_value::<Task>(payload.clone()).ok();
if let Some(task) = &task {
settle_task(ctx, task, Settled::Failed).await;
}
let notify_chat_id = payload.get("notify_chat_id").and_then(|v| v.as_i64());
let notify_message_id = payload.get("notify_message_id").and_then(|v| v.as_i64());
@@ -360,7 +431,7 @@ pub(crate) async fn dead_letter_notify(
ctx.sender,
notify_chat_id,
notify_message_id,
&format!("Task failed after retries: {message}"),
&failure_text(task.as_ref(), &message),
)
.await;
}
+2 -2
View File
@@ -17,8 +17,8 @@ pub struct ChatData {
pub edit_message: HashMap<i64, EditMessage>,
/// name -> HTML template containing "[]"
pub template: HashMap<String, String>,
/// site name (twitter/bsky/misskey/pixiv) -> user-supplied caption format
/// with {url} {author} {author_url} {title} {tags} placeholders.
/// site name (twitter/bsky/misskey/pixiv/bilibili) -> user-supplied caption format
/// with {url} {author} {author_url} {title} {content} {tags} placeholders.
pub message_format: HashMap<String, String>,
}