Compare commits

...
47 Commits
Author SHA1 Message Date
YoursFunny fb601f4d5d chore: bump version to 1.7.0 2026-09-18 01:17:45 +08:00
YoursFunny d8dd4fa91e test(cache): pin that pre-split entries still parse
A payload written before the title/content split has no `content` field;
`#[serde(default)]` is what keeps it readable, and the cache deletes any
payload it cannot parse — so dropping that default would silently evict
entries rather than degrade them. The test inserts the literal pre-split
JSON and asserts it comes back with its text left in `title` (no
migration: the entry lives one TTL and moving the text would only
reshuffle `/set_format` placeholders until it expires) and its stored
caption untouched.
2026-09-18 01:14:09 +08:00
YoursFunny af96caff40 feat(send): quote a long post's text in an expandable blockquote
A post whose text (the split `title` plus `content`, joined by
`site::compose_text`) reaches `CAPTION_QUOTE_TEXT_CHARS` — default 200,
`0` disables — now has that text wrapped in `<blockquote expandable>`
inside its caption, leaving the URL and author line outside the quote.

Applied at the send boundary (`send_media_sequence`, `send_animation` and
the inline answers), where the caption is already truncated and the same
cache snapshot supplies the text, so a fresh send, a link-cache resend
and a queued retry all decide identically. The text is located as what
follows the author link, with the visible prefix accepted as a match
because `truncate_caption` may cut inside it — that keeps the longest
posts, the ones that most need folding, quoted. Captions whose layout
moves the text elsewhere (pixiv's title-inside-a-link, a `/set_format`
that puts `{title}`/`{content}` first) stay unquoted rather than risking
a blockquote nested in a tag, and a caption that already carries one is
never wrapped again.

Telegram measures a caption *after entities parsing*, so the tags cost no
length and the 1024-character limit cannot be breached; retries replay
the unwrapped caption, so a threshold change takes effect immediately.
The edit-before-forward rewrite stays unquoted by design.

Verified against Telegram: a media-group caption built this way comes
back with `caption_entities` `url` @0, `text_link` @50,
`expandable_blockquote` @56 — the quote starts after the author line.
2026-09-18 00:50:30 +08:00
YoursFunny 52184ba6fb refactor(x-media): split the post title from its content
`Fetched.title` carried whatever text the platform had — a tweet's body,
a bilibili dynamic's body, a pixiv artwork's title — which was enough
while x/twitter (no title at all) set the shape. The platforms actually
disagree: pixiv has a title *and* a description, bilibili has an opus
headline *and* a body. Posts now carry both:

- `title`: the platform's title (a pixiv artwork title, a bilibili opus
  headline or video card title), empty on text-only platforms;
- `content`: the body (tweet / bsky / misskey text, bilibili dynamic
  body, and pixiv's description — fetched for the first time here and
  flattened from the app API's HTML to plain text).

`{content}` joins the caption-format placeholders, so a custom
`/set_format` can include a pixiv description. The built-in captions keep
producing byte-identical output: `compose_text` joins the two fields the
same way the single field already was, and bilibili's forward marker
(`//@author:`) now lands in `content` behind the head line's `title`.

`CachedPost.content` is `#[serde(default)]`, so link-cache entries and
queued task payloads written before the split still parse, their text
living in `title`.
2026-09-18 00:44:49 +08:00
YoursFunny 0eb4e5c78d fix(sites): request the opus serialization so bilibili posts keep their text
An image/text post fetched without `features=itemOpusStyle` comes back in
bilibili's legacy shape, where the post's body and headline are gone
completely — `desc: null`, no `major.opus` — so `title` (and `{title}`)
stayed empty for exactly the posts that do have content
(`opus/1248857553488576532`: legacy `desc: null`, flagged
`major.opus.summary.text = "[doge_金箍]黑白搭配"`). The same flag also
moves the pictures to `major.opus.pics` (key `url`, not `src`).

Text now falls back opus (headline + body) → `desc.text` → archive card
title; media falls back `opus.pics` → `draw.items` → archive cover, so
the legacy shapes keep working if the flag is ever retired.

Verified live: the reported link now yields title "[doge_金箍]黑白搭配"
with its picture; AV dynamics keep their card title; forwards keep
`desc.text` and gain `//@` composition unchanged.
2026-09-17 22:34:11 +08:00
YoursFunny 5c51de217a fix(sites): use the archive title when a bilibili dynamic has no body
A 视频投稿动态 (`DYNAMIC_TYPE_AV`) carries no body at all: the API
answers `desc: null` and the content is the archive card, so `title`
(and `{title}` in caption formats) stayed empty for the most common
dynamic type. Audited 24 live dynamics: every dynamic that *has* text
(a 图文 post, a forward, a text post) keeps it in
`module_dynamic.desc.text` — only the AV card has none, so the video
title now stands in, mirroring pixiv whose `title` is the artwork title
rather than post text.

Also records two API observations in comments/docs: an id that cannot
exist answers `4101105 请求数据发生错误` (kept on the permanent arm), and
the feed endpoints strip `desc.text` so only the detail endpoint shows
whether a post has text.
2026-09-17 22:12:53 +08:00
YoursFunny c1f5d3ca54 feat(sites): add bilibili dynamic support (images and animated images)
Fetch `t.bilibili.com/<id>`, `www.bilibili.com/opus/<id>`,
`t.bilibili.com/h5/dynamic/detail/<id>` and `m.bilibili.com/dynamic/<id>`
through the anonymous `/x/polymer/web-dynamic/v1/detail` endpoint (no
cookie, no WBI signature; the site adds the device cookies
`/x/frontend/finger/spi` hands out, which is what lifts bilibili's
`-352` risk control).

Media: the `major.draw` grid (`.gif` sources become animations, the rest
photos with a downscaled `@518w.jpg` thumbnail used both as preview and
as the oversized fallback), an attached video's cover, and the quoted
dynamic's media for forwards. The video stream itself is not resolved;
`b23.tv` short links stay unmatched (they mostly point at videos, so
matching them would turn a silently ignored link into a failure reply).

`-352`/`-412` map to a retryable error so the queue backs off instead of
dropping the post; a removed dynamic (`500`) is permanent.

Registry-driven, so no bot-side code changes beyond the site lists in the
command replies; found while researching nazurin and
telegram-bili-feed-helper (see BILIBILI_PLAN.md).
2026-09-17 20:48:15 +08:00
YoursFunny 1bb6968108 Merge pull request #3 from TheFunny/dependabot/cargo/rand-0.10.2
build(deps): bump rand from 0.8.8 to 0.10.2
2026-09-17 17:05:30 +08:00
YoursFunny 14b444d109 Merge master into the rand 0.10 migration
Resolve the manifest and lockfile overlap with the rusqlite 0.40 and zip 8.6
bumps that landed on master after this branch was opened:
- crates/xmedia-bot/Cargo.toml: keep rusqlite 0.40 and rand 0.10
- Cargo.lock: re-point x-media/xmedia-bot at rand 0.10.2; teloxide keeps its
  own rand 0.8.8 (it requires ^0.8.5), and cargo pruned the now-unused
  version_check entry

Verified on the merged tree: cargo metadata --locked, cargo fmt --check,
cargo clippy --workspace --all-targets --locked -- -D warnings,
cargo test --workspace --locked (69 + 73 tests).
2026-09-17 16:59:27 +08:00
YoursFunny de3105d4cd Merge pull request #2 from TheFunny/dependabot/cargo/zip-8.6.0
build(deps): bump zip from 2.4.2 to 8.6.0
2026-09-17 16:54:44 +08:00
YoursFunny c46103a23f Merge pull request #1 from TheFunny/dependabot/cargo/rusqlite-0.40.2
build(deps): bump rusqlite from 0.32.1 to 0.40.2
2026-09-17 16:54:00 +08:00
YoursFunny 9529f64b41 fix(send): migrate rand usage to the 0.10 API
rand 0.10 renamed the entry points dependabot's version bump alone cannot
follow: `thread_rng()` -> `rng()`, `gen_range()` -> `random_range()` and the
`Rng` trait -> `RngExt`.

The jitter only needs one float, so use the free function
`rand::random_range(0.2..0.8)` and drop the now-unused trait import instead
of importing `RngExt`.

Verified: cargo fmt --check, cargo clippy --workspace --all-targets -- -D
warnings, cargo test --workspace --locked (retry_delay_seconds_bounds keeps
the 0.2..0.8 jitter window).
2026-09-17 16:43:01 +08:00
dependabot[bot] d0de17329d build(deps): bump rand from 0.8.8 to 0.10.2
Bumps [rand](https://github.com/rust-random/rand) from 0.8.8 to 0.10.2.
- [Release notes](https://github.com/rust-random/rand/releases)
- [Changelog](https://github.com/rust-random/rand/blob/master/CHANGELOG.md)
- [Commits](https://github.com/rust-random/rand/compare/0.8.8...0.10.2)

---
updated-dependencies:
- dependency-name: rand
  dependency-version: 0.10.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-17 03:43:17 +00:00
dependabot[bot] 9af37e92b4 build(deps): bump zip from 2.4.2 to 8.6.0
Bumps [zip](https://github.com/zip-rs/zip2) from 2.4.2 to 8.6.0.
- [Release notes](https://github.com/zip-rs/zip2/releases)
- [Changelog](https://github.com/zip-rs/zip2/blob/master/CHANGELOG.md)
- [Commits](https://github.com/zip-rs/zip2/compare/v2.4.2...v8.6.0)

---
updated-dependencies:
- dependency-name: zip
  dependency-version: 8.6.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-17 03:43:10 +00:00
dependabot[bot] e68a1dbd30 build(deps): bump rusqlite from 0.32.1 to 0.40.2
Bumps [rusqlite](https://github.com/rusqlite/rusqlite) from 0.32.1 to 0.40.2.
- [Release notes](https://github.com/rusqlite/rusqlite/releases)
- [Changelog](https://github.com/rusqlite/rusqlite/blob/master/Changelog.md)
- [Commits](https://github.com/rusqlite/rusqlite/compare/v0.32.1...v0.40.2)

---
updated-dependencies:
- dependency-name: rusqlite
  dependency-version: 0.40.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-17 03:43:05 +00:00
YoursFunny b9c6d16ff0 docs(ci): note why the docker job skips actions/checkout
build-push-action defaults to the Git context, so BuildKit clones the
repo itself and the job never needs the workspace.
2026-09-17 11:38:49 +08:00
YoursFunny ec65c3ce74 chore: bump version to 1.6.0
Bump both crates (x-media, xmedia-bot) and refresh the lockfile to the
latest semver-compatible releases. No direct dependency or manifest
requirement changed.
2026-09-17 11:21:58 +08:00
YoursFunny 893ab7a1e0 feat(commands): /test sends the media, /debug takes over the parse report
- `/test <url>` now runs the ordinary link pipeline and actually sends the
  media, but with the chat's post-send actions suppressed: no channel forward,
  no edit-before-forward prompt. It is the same code path as a normal link
  (same caption/format handling, link cache, retries, dead-letter
  notification), so "does this link work?" is answered by the send itself.
- `/debug <url>` keeps what `/test` used to do: fetch and reply with the HTML
  parse report, sending/caching/forwarding nothing.
- `urls::url_media` takes a `PostSend` mode (`FromChat` for the URL workers,
  `Suppressed` for `/test`); `build_send_task` maps it to the task's
  `edit_before_forward`/`forward_channel_id`. Notification ids stay set in both
  modes, so a queued retry still reports a dead-letter to the chat.
- `/test` rejects an unsupported URL with the same message the old parse-only
  command used (the URL flow would otherwise ignore it silently).
- Report builder renamed `test_parse_report` -> `debug_report` (with the cap
  constant), `parse_test_arg` -> `parse_arg_remainder` (now shared by both
  commands). README/README.en command tables and AGENTS.md updated; `/help`
  descriptions come from the enum.

Tests: +3 (normal flow still honours the chat's settings, `/test` sends with
them suppressed and keeps the cache entry, `build_send_task` mode mapping). The
suppression test was verified to fail when the mode is ignored.
fmt/clippy clean, 73 + 69 tests pass.
2026-09-17 02:25:42 +08:00
YoursFunny bd032e3d68 ci: lock the dependency set, verify the build inputs on PRs, harden the jobs
- `--locked` on every cargo invocation (ci.yml clippy/test/build, both
  Dockerfile builds). The version bump edits Cargo.lock by hand, so a stale
  lock must fail loudly instead of being silently re-resolved: CI would
  otherwise test a different dependency set than the one committed — and than
  the one the released image is built from.
- docker.yml: build the image (no push, no registry login, read-only build
  cache) on pull requests touching the build inputs. The Dockerfile's
  stub-source machinery, the ffmpeg download and the entrypoint previously
  only ran at release time. Also: a release tag must equal both crate versions
  before anything is built (the binary carries no version, so `v1.5.1` with
  manifests at 1.5.0 used to publish silently wrong tags), `FFMPEG_URL`/
  `FFMPEG_SHA256` are taken from repository variables when set, and the
  unused `setup-qemu-action` step is gone (single-arch build; the comment says
  what arm64 would need).
- ci.yml: `concurrency` cancels superseded runs, `permissions: contents: read`,
  `RUST_BACKTRACE=1`, job timeouts, and a release-profile build of the same
  package the Dockerfile builds (the profile was otherwise never compiled
  before a merge). The `live` job narrows to `-p x-media`: every network- or
  secret-gated test lives there, and the bot crate's offline suite already ran
  in the `test` job. Timeout is 45 min because the release build is cold on
  the first run — a timeout there would kill the job before rust-cache could
  save its cache, leaving every later run cold too.
- Actions pinned to commit SHAs (Dependabot keeps them current);
  `dtolnay/rust-toolchain` stays on its channel ref by design.
- .github/dependabot.yml: crates (patch bumps grouped), action pins, Docker
  base images — the audit gate reports advisories, this is what moves them.
- tokio's `sync` feature is now declared instead of arriving transitively via
  teloxide; `.dockerignore` drops docs and markdown.

Verified locally: `cargo fmt --check`, `cargo clippy --workspace
--all-targets --locked`, `cargo test --workspace --locked` (70 + 69 pass),
`cargo build --release --locked` (6m03s cold, the 15.9 MB stripped binary
starts and registers 10 commands), the tag/version gate against both a
matching and a mismatching tag, and YAML parsing of all three workflow files.
2026-09-17 02:01:43 +08:00
YoursFunny dbda6ec1c2 refactor(send): split the 2100-line module by concern
send.rs had grown back into the shape handlers.rs was split out of: payload
types, error classification, the download-and-reupload fallback, the senders
and the whole post-send/queue shell in one file. Split by concern, leaving
call sites (`crate::send::x`) unchanged:

- `send/input_media.rs`: payload → `InputFile`/`InputMedia` selection and
  `build_media_group` (with its caption-on-first-item rule).
- `send/upload.rs`: the fallback pipeline (download with the upload cap,
  photo downscale handoff, smaller-URL fallback, multipart upload).
- `send/post_send.rs`: link-cache write, the `KEEP_ALIVE` registry for locally
  produced media, `settle_task`, the post-send actions and the queue entry
  points; the parts other modules call are re-exported.
- `send/mod.rs`: payloads, error classification, classification helpers and
  the senders themselves, plus the test module.

No behaviour change: 128 + 1326 + 366 + 345 lines, 70 tests still pass.
AGENTS.md updated for the new layout and for `ctx.rs`.
2026-09-17 01:41:28 +08:00
YoursFunny 0a9ff58a69 refactor: cover the edit/answer surface in MediaSender, test the button flows
docs/architecture-refactor.md §3 sketched the trait with "按需扩展:
edit_message_caption / delete_message / answer_callback_query …", but only the
five send methods landed, so `callback.rs` and the edit-before-forward caption
swap were stuck on the concrete `Bot` and remained untested (AGENTS.md still
lists callback.rs as untestable).

- `MediaSender` gains `answer_callback_query`, `edit_message_caption` (HTML
  parse mode baked in, every caller uses it) and `delete_message`; the mock
  records call order plus the texts, captions and answer toasts, so tests can
  assert what the user saw.
- `send_message` now returns the sent message id instead of the whole
  `Message`: the only consumer of the value is the edit-before-forward prompt
  (which keys its record by it), and returning a `Message` forced every mock
  to build a teloxide type. `reply`/`reply_html` follow.
- `callback.rs`: the dptree entry only unpacks the update; `handle_callback`
  takes plain values + `&AppContext`. `handlers/mod.rs::edit_message_handler`
  likewise takes the values the reply carries. Admin/setup APIs
  (`get_chat`, `get_chat_administrators`, `get_me`, `set_my_commands`) stay on
  the concrete `Bot`: they are not user flows worth a trait.
- The scripted mock moves to `parking_lot::Mutex` (no poisoning unwraps).

Tests: +11 (template button, forward ok/no-channel/retryable, expired+unknown
prompt, caption swap via template, escaping of user text into the caption,
failed swap still consuming the reply, prompt record written by post_send).
fmt/clippy clean, 70 + 69 tests pass.
2026-09-17 01:31:30 +08:00
YoursFunny c2d7c8406e refactor: finish the phase-B seam for the post-send path, funnel settlement
docs/architecture-refactor.md §3 stopped half-done: `url_media` got an injected
`AppContext`, but `send.rs`'s post-send half kept reaching for the process-wide
`CHAT_STORE`/`TASK_QUEUE`/`LINK_CACHE` statics, so the whole shell after a
successful send (edit-before-forward prompt, channel forward, retry enqueue,
cache write) had no test and no way to get one.

- `ctx.rs` now owns `AppContext` (sender + the three stores + config) with
  `from_statics` for production and a `CONTEXT` static for the spawned worker
  closures; `handlers/urls.rs` drops its private copy and the duplicated
  assembler, and the queue handler/dead-letter callbacks take the context
  (main wires them with `CONTEXT`).
- `send_media_sequence`/`send_animation`/`forward_messages`/`post_send_actions`
  take `&AppContext`; the cache write goes through the injected cache.
- New `settle_task(ctx, task, Sent|Failed)` is the single place that ends a
  task: release its keep-alive temp media, and drop the link-cache entry only
  on failure. All five former call sites funnel through it — the earlier
  keep-alive leak existed precisely because one of them had to remember.
  `invalidate_cache`/`invalidate_cache_with` (static + injected pair, the
  latter only existing because of the former) collapse into one private fn.
- `ctx::test_support::TestStores` gives tests a tempdir store set + context;
  `handlers/urls.rs` tests use it instead of hand-rolled setup.

Tests: +5 (post-send forward ok / queued / notified, settle Sent/Failed); the
post-send and settle paths were previously untested. fmt/clippy clean,
60 + 69 tests pass.
2026-09-17 01:27:05 +08:00
YoursFunny abdc27ed5e docs: resync AGENTS.md and the doc comments with the code
AGENTS.md:
- db.rs row claimed a per-store connection pool; there is one shared pool for
  all three tables (statics.rs builds it once).
- retry enqueue moved to send.rs, noted in both handlers rows.
- queue row now names both notifies (workers' + the sweep's).
- Retries bullet documents fetch vs fetch_once.
- test count ~125 -> ~135, the untested-files list no longer claims state.rs
  and handlers.rs are untested, and the live-test inventory mentions the
  token-gated, not-#[ignore]d pixiv download test that makes a local
  `cargo test --workspace` hit the network.
- /bot_dict is admin-only now.

Code docs:
- site/mod.rs: the module doc pointed new sites at `fetch_once` (a name that
  did not exist then and now means a single-attempt fetch) -> `SITES`; the
  cache_key/SITES/Site docs still said "twitter -> bsky -> pixiv" (misskey
  is registered third); RenderData now documents which fields are escaped
  and why url/author_url are not.
- state.rs, callback.rs: drop the pre-misskey site list and the `<name>`
  that rustdoc read as an HTML tag.
- Fixed the remaining rustdoc links/warnings: `cargo doc --workspace
  --no-deps` is now warning-free (was 8).
- docs/site-registry-refactor.md: §1 describes the pre-refactor state; said so.

No behavior change. fmt/clippy clean, 55 + 68 tests pass (the live pixiv
download test flaked on a CDN body timeout, as before).
2026-09-16 23:28:46 +08:00
YoursFunny 3fb8421c3a perf: stop the sweep stealing worker wakeups, retry-free inline fetch
- queue: the lease-expiry sweep waited on the workers' `Notify`. `notify_one`
  stores a permit, so a sweep wakeup could consume the one meant for a worker,
  which then blocked on `notified()` (it only waits when the table looked
  empty, i.e. indefinitely) with a due row sitting there. The sweep now has
  its own notify, woken only by stop.
- x-media: split fetch's retry loop into `fetch` (3 attempts, unchanged) and
  `fetch_once` (1 attempt); inline queries use the latter — the 800ms debounce
  plus 1s/2s backoffs were outlasting the answer window of the query.
- send: chunk_media_items now moves items out of the input Vec instead of
  requiring `T: Clone` and copying every payload.

fmt/clippy clean, 55 + 69 tests pass.
2026-09-16 21:23:53 +08:00
YoursFunny 475cfd18f9 fix: harden the debug command, link cache and rate limiter
- commands: /bot_dict dumped the whole chat state to any member of the chat
  and could exceed Telegram's 4096-char message limit (the send then failed
  and bubbled up as a handler error). It is now admin-only and capped at
  MAX_DEBUG_DUMP_CHARS; README, README.en and the /help description updated.
- send: the edit-before-forward template buttons were built from a HashMap
  walk, so their order changed between prompts. Now sorted by name.
- link_cache: an unparseable payload (older schema) was reported as a miss
  but left in place, re-failing the parse on every later hit; the row is
  dropped on read.
- handlers: a link handed to the URL workers after the channel closed
  (shutdown) was discarded silently; it is now logged.
- rate_limit: LIMITERS kept one bucket per chat that ever sent media,
  forever. The periodic sweep now drops buckets that are idle (refilled to
  capacity) and not held by an in-flight sender; acquire's refill was
  factored into a shared helper used by the idle check.

Tests: +3 (corrupted row dropped, sorted markup, idle-bucket pruning); the
cache one was verified to fail before the fix. fmt/clippy clean, 55 + 69.
2026-09-16 21:16:14 +08:00
YoursFunny 0a577600fd fix: repair four correctness defects in the send/state/handler paths
- send: the PHOTO_INVALID_DIMENSIONS marker never matched (the description is
  lower-cased, the marker was not), so oversized photos sent by URL were
  classified Permanent instead of taking the download-and-downscale fallback.
- inline: the debounce state was one global slot, so a second user's query
  cancelled the first user's pending answer entirely; it is now per user.
- state: prune_expired wrote back a stale snapshot without the per-chat lock,
  clobbering a concurrent update() (lost edit-message record -> "Expired");
  it now re-reads and prunes under the same lock update() uses.
- send/queue: a task dead-lettered on retry exhaustion kept its keep-alive
  temp media (ugoira MP4) alive until process exit; dead_letter_notify now
  releases it, and enqueue_retry releases when the enqueue itself fails.

Also folds the duplicated retry enqueue in post_send_actions into
send::enqueue_retry (single clock source, single place that releases).

Tests: +6 (marker, per-user debounce x3, keep-alive release, prune contract
x2); the marker and keep-alive cases were verified to fail before the fix.
cargo fmt/clippy clean, 52 + 69 tests pass.
2026-09-16 21:11:39 +08:00
YoursFunny 32254fa807 chore: bump version to 1.5.0 2026-09-07 21:26:24 +08:00
YoursFunny 11c04b66dc feat: add Misskey (misskey.io) fetch support
Fourth site adapter: POST /api/notes/show, renote-aware caption and
media normalization, DriveFile type → Illustration/Animated/Video.
Empty thumbnailUrl strings filtered out in thumbnail_for.
2026-09-07 21:25:52 +08:00
YoursFunny 2f741e5f4b refactor(send): apply ponytail audit cuts 2, 4, 6
- updated_sequence_task: clone the Task and mutate the two fields
  instead of rebuilding all 12 by hand (-22 lines; new fields no
  longer need a sync here)
- unify unix_now with db::now_f64 (unix_now() = now_f64() as i64),
  moved to db.rs next to its clock source
- classify_to_send_error takes the MediaFetchFailure label, folding
  the duplicated inline match in send_batch_via_upload (-8 lines)
2026-09-07 19:19:56 +08:00
YoursFunny 89c4642e1c fix(lint): resolve clippy warnings from rust 1.98
- photo.rs: chunks_exact(4)/(2) -> as_chunks::<N>().0
  (chunks_exact_to_as_chunks, the new lint prefers the
  compile-time-checked slice split)
- send.rs: box the Task inside SendError so the error fits the
  result_large_err limit (Task is ~400 bytes; the error now moves
  through Result as a pointer); unbox with *task at the two
  enqueue_retry call sites (handlers/urls.rs, handlers/callback.rs)

cargo clippy --workspace --all-targets is now warning-free; the
remaining proc-macro-error2 future-incompat note is upstream
(teloxide -> aquamarine) and unfixable locally. Full test suite passes.
2026-09-07 17:00:38 +08:00
YoursFunny 3f6a0f034a chore: bump version to 1.4.0
1.3.0 → 1.4.0: new features (DATA_DIR config, /test blockquote HTML
report) plus the twitter entity-decode, post-send-actions and queue
lease-heartbeat fixes.
2026-08-16 17:43:54 +08:00
YoursFunny 90a011e978 feat(commands): wrap the /test caption in a blockquote (HTML report)
Replaces the strip-tags plain-text rendering: the /test reply is now an
HTML message (reply_html helper with ParseMode::Html). Raw fields (url,
source_url, title, author_url, media urls) are escaped, the pre-escaped
render fields are embedded as-is, and the caption is wrapped in
<blockquote>...</blockquote> so the report shows it exactly as it will
render in the sent media caption — escaped text and clickable links
included, no literal &amp;/&lt;/&gt; and no raw markup.
2026-08-16 17:31:54 +08:00
YoursFunny 12a065846c feat(statics): make the SQLite path configurable via DATA_DIR
The DB file was hardcoded to CWD-relative data/task_queue.db — a footgun
for systemd/cron deployments and a confusing startup failure when the
data/ dir did not exist (SQLite never creates parent dirs).

db_path() now reads DATA_DIR (default data, CWD-relative, unchanged for
local runs and the docker-compose ./data mount) and creates the
directory automatically. README/README.en.md env tables and AGENTS.md
document the new variable.
2026-08-16 17:07:33 +08:00
YoursFunny dca1eff1c9 ci: add a cargo-audit dependency vulnerability gate
Runs actions-rust-lang/audit after the offline tests in the test job: a
crate in Cargo.lock with an unfixed security advisory fails the build.
Verified locally against the current lockfile (0 vulnerabilities; the 3
warnings — unmaintained dotenv/proc-macro-error2 and transitive anyhow
unsoundness — do not fail by default).
2026-08-16 17:06:28 +08:00
YoursFunny 894a9ebf4a docs: correct the user-facing string language claim in AGENTS.md
AGENTS.md claimed user-facing bot strings are Chinese, but every
reply/send_message string in the code is English (Hello!, Send failed,
No media found, Reply to edit message, ...). README stays Chinese;
update both the overview line and the convention line to state the
actual split.
2026-08-16 17:02:05 +08:00
YoursFunny c968891ff6 fix(commands): render the /test caption as plain text
The report's caption line still showed the raw HTML markup
(<a href="...">...</a>). strip_html_tags now drops the tags (keeping
the visible text; the links are already reported via source_url /
author_url) and the remaining entity-encoded text is decoded — the
strip runs on the escaped caption so a tweet text like >^ω^< survives
instead of being eaten as markup. Custom-format captions contain no
tags and pass through unchanged.
2026-08-16 17:01:47 +08:00
YoursFunny 0087bd01ac fix(queue): heartbeat the lease so long tasks are not re-processed
The lease was set once to now + LOCK_TTL_SECONDS (120 s) with no
renewal. Tasks that legitimately take longer — slow CDN downloads,
ugoira encodes, rate-limited batch forwards (a 100-message channel copy
waits ~4 min on the per-chat token bucket) — had their lease expire
mid-run; the 30 s expiry sweep flipped the row back to pending and
another worker processed it again, double-sending.

run_with_lease now drives the handler through tokio::select! and
refreshes locked_until every 30 s while it runs. The heartbeat lives in
the same future as the handler, so a panicking worker still lets the
sweep recover the row (no leaked task keeping the lease fresh forever).
2026-08-16 17:00:38 +08:00
YoursFunny 4cb40909c5 fix(send): run post_send_actions after retried sends
A task only reaches the queue after a failed send, so the fresh attempt
never ran post_send_actions (edit-before-forward prompt / channel
forward) — it failed before that point. The old guard skipped
post_send_actions for resumed tasks (batch_index > 0 or sent ids
present), which meant any send that needed a retry after partial
progress silently lost its forward and edit prompt.

post_send_actions is now run unconditionally on a successful queue send;
it executes exactly once, after the whole sequence completed.
2026-08-16 17:00:31 +08:00
YoursFunny f260f41755 docs: align AGENTS.md and READMEs with the current code
AGENTS.md: document the twitter API entity decode and the /test report
HTML-decoded display; add the missing db.rs / media_sender.rs /
rate_limit.rs module rows; fix the statics location (handlers/statics.rs);
refresh test counts (~115, twitter live 5, pixiv api.rs 1, photo heavy
test); versioning convention now includes README.en.md.
README.md / README.en.md: add TELOXIDE_PROXY to the env variable list.
2026-08-16 16:31:22 +08:00
YoursFunny af901caddb fix(twitter): decode API HTML entities so captions escape exactly once
Twitter's syndication and GraphQL APIs return tweet text and display
names pre-escaped for HTML (&gt; &lt; &amp; &#39;); the caption builder
escaped the text again, so sent messages showed literal entities (e.g.
>^ω^< came back as &gt;^ω^&lt;). from_syndication_json now decodes the
API text before storing it — both the syndication path and the
TWITTER_AUTH_TOKEN GraphQL fallback route through it — so the caption
escapes exactly once and renders correctly.

The /test report is a plain-text message but printed the pre-escaped
caption and render fields; it now HTML-decodes them for display so the
report shows the rendered text.
2026-08-16 16:31:16 +08:00
YoursFunny c0af42b1cc chore: bump version to 1.3.0 2026-08-15 15:49:40 +08:00
YoursFunny d0810217b4 feat(commands): add /test debug command that reports link parse results only 2026-08-15 15:48:57 +08:00
YoursFunny e6ba178983 feat(send): add per-chat token bucket rate limiting 2026-08-15 00:34:59 +08:00
YoursFunny ae69d72930 refactor(handlers): inject AppContext into url_media; cover the full URL pipeline 2026-08-14 23:38:41 +08:00
YoursFunny 1e30815a10 docs: mark architecture refactor phases A and B implemented 2026-08-14 22:04:30 +08:00
YoursFunny 50206a9056 refactor(send): introduce MediaSender seam; add mock-based fallback tests 2026-08-14 22:04:16 +08:00
YoursFunny c9e72fda70 refactor(handlers): split monolithic handlers.rs into modules 2026-08-14 21:50:40 +08:00
48 changed files with 8557 additions and 3654 deletions
+4
View File
@@ -26,6 +26,10 @@
*.db
LICENSE
README.md
# Documentation and scratch files: the build only ever reads the manifests,
# `crates/` and the entrypoint script.
docs/
*.md
data/
cert/
nginx-certs/
+32
View File
@@ -0,0 +1,32 @@
version: 2
# Pairs with the `actions-rust-lang/audit` gate in ci.yml: the gate reports
# advisories in Cargo.lock, this is what actually moves the dependencies.
# Patch bumps are batched into one PR; minor/major stay separate so they get
# reviewed and tested individually.
updates:
- package-ecosystem: cargo
directory: /
schedule:
interval: weekly
open-pull-requests-limit: 5
groups:
cargo-patch:
applies-to: version-updates
patterns: ['*']
update-types: ['patch']
# The workflow actions are pinned to commit SHAs; that pin is what makes
# bumping them a manual chore, so let the bot do it.
- package-ecosystem: github-actions
directory: /
schedule:
interval: weekly
open-pull-requests-limit: 5
# The Dockerfile's base images (rust:1-bookworm, debian:bookworm-slim).
- package-ecosystem: docker
directory: /
schedule:
interval: monthly
open-pull-requests-limit: 3
+54 -12
View File
@@ -4,14 +4,19 @@ name: CI
# job that exercises the real source sites and the token-gated pixiv tests.
#
# Layering:
# test — fmt + clippy + the full offline unit suite. Runs on every push
# and PR, including forks (it needs no secrets).
# test — fmt + clippy + the full offline unit suite + a release-profile
# build + cargo-audit dependency gate. Runs on every push and PR,
# including forks (it needs no secrets).
# live — the #[ignore]d live-network tests plus the pixiv tests that are
# gated on PIXIV_REFRESH_TOKEN. Runs on schedule / manual dispatch
# / tag pushes only, because pull requests from forks cannot read
# repository secrets. continue-on-error keeps a flaky external site
# from blocking, while the run still records the outcome.
#
# Every action is pinned to a commit SHA (Dependabot keeps the pins current);
# `dtolnay/rust-toolchain` deliberately stays on its channel ref, because the
# ref itself is what selects the toolchain (`@stable` = install stable).
#
# Test gating convention (keep in sync with AGENTS.md "Testing & QA"):
# - pure unit tests: plain #[test] / #[tokio::test], always run.
# - live-network tests: #[ignore = "live network: ..."], only run here.
@@ -27,38 +32,75 @@ on:
- cron: '0 3 * * 1'
workflow_dispatch:
permissions:
contents: read
# A newer push to the same ref supersedes the older run; without this every
# intermediate commit of a PR branch keeps a runner busy to completion.
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
env:
# Panicking tests print their backtrace; free when nothing fails.
RUST_BACKTRACE: 1
jobs:
test:
runs-on: ubuntu-latest
# Generous on purpose: the release-profile build below is cold on the very
# first run (thin LTO + codegen-units = 1 across every dependency), and a
# timeout there would kill the job *before* rust-cache saves its cache —
# leaving every later run cold again.
timeout-minutes: 45
steps:
- uses: actions/checkout@v7
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: dtolnay/rust-toolchain@stable
with:
components: clippy, rustfmt
- uses: Swatinem/rust-cache@v2
- uses: Swatinem/rust-cache@6323deb102c322ba6fcbdcafc7e3dddab59af2b6 # v2
# `--locked` on every cargo invocation: the version bump edits
# Cargo.lock by hand (AGENTS.md), so a stale lock must fail here instead
# of being silently re-resolved — otherwise CI tests a different
# dependency set than the one committed, and than the one the released
# image is built from.
- name: Check formatting
run: cargo fmt --check
- name: Lint (deny warnings)
run: cargo clippy --workspace --all-targets -- -D warnings
run: cargo clippy --workspace --all-targets --locked -- -D warnings
- name: Run offline tests
run: cargo test --workspace
run: cargo test --workspace --locked
# The release profile (lto/strip/codegen-units=1, overflow checks off)
# was otherwise only exercised by the Docker build on master/tag. Same
# package the Dockerfile builds; the cache keeps it cheap after the
# first run.
- name: Build release profile
run: cargo build --release --locked -p xmedia-bot
# Dependency vulnerability gate: fails the build when a crate in
# Cargo.lock has an unfixed security advisory. Unmaintained/unsound
# *warnings* (dotenv, proc-macro-error2, anyhow transitive) do not fail
# the build by default; the advisory DB is cached across runs.
- name: Audit dependencies
uses: actions-rust-lang/audit@72c09e02f132669d52284a3323acdb503cfc1a24 # v1
live:
needs: test
if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch' || startsWith(github.ref, 'refs/tags/v')
runs-on: ubuntu-latest
timeout-minutes: 30
continue-on-error: true
env:
PIXIV_REFRESH_TOKEN: ${{ secrets.PIXIV_REFRESH_TOKEN }}
TWITTER_AUTH_TOKEN: ${{ secrets.TWITTER_AUTH_TOKEN }}
steps:
- uses: actions/checkout@v7
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
- uses: dtolnay/rust-toolchain@stable
- uses: Swatinem/rust-cache@v2
# Full suite: with the secret present, the pixiv token-gated tests run;
# without it they skip themselves. Live tests stay #[ignore]d here.
- uses: Swatinem/rust-cache@6323deb102c322ba6fcbdcafc7e3dddab59af2b6 # v2
# Everything network- or secret-gated lives in x-media, and the bot
# crate's suite (MockSender + tempdir stores, no network) already ran in
# the `test` job — rebuilding it here bought nothing.
- name: Run token-gated tests
run: cargo test --workspace
run: cargo test -p x-media --locked
# The live-network tests, by the "live" name filter (all #[ignore]d).
- name: Run live-network tests
run: cargo test --workspace -- --ignored live
run: cargo test -p x-media --locked -- --ignored live
+75 -11
View File
@@ -1,28 +1,75 @@
name: Build Docker Image
# Release builds (master / v* tags) plus a build-only check on pull requests
# that touch anything the image depends on — the Dockerfile's stub-source
# machinery, the ffmpeg download and the entrypoint are exactly the parts that
# would otherwise break only at release time.
#
# Actions are pinned to commit SHAs (Dependabot keeps the pins current).
on:
push:
tags:
- v*
branches:
- master
pull_request:
paths:
- Dockerfile
- docker-entrypoint.sh
- .dockerignore
- Cargo.toml
- Cargo.lock
- .github/workflows/docker.yml
- 'crates/**/Cargo.toml'
env:
APP_NAME: telegram-twitter-media-bot
DOCKERHUB_REPO: yoursfunny/telegram-twitter-media-bot
permissions:
contents: read
# Serialize runs per ref. Never cancel in progress: a killed run would drop a
# half-finished image push.
concurrency:
group: docker-${{ github.ref }}
cancel-in-progress: false
jobs:
# A tag push and a branch push to the same commit fire two workflow runs;
# build only once. Tag runs always build; master runs build only when the
# pushed commit is not already tagged (the tag run covers it).
should-build:
runs-on: ubuntu-latest
timeout-minutes: 10
outputs:
build: ${{ steps.check.outputs.build }}
steps:
- uses: actions/checkout@v7
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
with:
fetch-depth: 0
# A release tag is the version claim: the manifests are bumped by hand,
# so `v1.5.1` with `Cargo.toml` still at 1.5.0 would publish an image
# whose tag lies about what is inside it (the binary carries no version).
- name: Verify the tag matches both crate versions
if: startsWith(github.ref, 'refs/tags/v')
shell: bash
run: |
tag="${GITHUB_REF_NAME#v}"
status=0
for manifest in crates/x-media/Cargo.toml crates/xmedia-bot/Cargo.toml; do
# tr -d '\r': a CRLF checkout (core.autocrlf on Windows) would
# otherwise yield "1.5.0\r" and false-fail every tag.
version="$(sed -n 's/^version = "\(.*\)"/\1/p' "$manifest" | head -1 | tr -d '\r')"
if [ "$version" != "$tag" ]; then
echo "::error file=$manifest::$manifest is at $version but the tag is v$tag"
status=1
else
echo "$manifest: $version matches v$tag"
fi
done
exit "$status"
- id: check
shell: bash
run: |
@@ -33,14 +80,20 @@ jobs:
echo "build=true" >> "$GITHUB_OUTPUT"
fi
# No `actions/checkout` here on purpose: `docker/build-push-action` defaults
# to the Git context (`https://github.com/<owner>/<repo>.git#<ref>`), so
# BuildKit clones the repo itself and authenticates with the automatic
# github.token. Adding `context: .` below without a checkout step would hand
# BuildKit an empty workspace.
docker:
needs: should-build
if: needs.should-build.outputs.build == 'true'
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- name: Docker meta
id: meta
uses: docker/metadata-action@v6
uses: docker/metadata-action@dc802804100637a589fabce1cb79ff13a1411302 # v6
with:
images: ${{ env.DOCKERHUB_REPO }}
tags: |
@@ -49,15 +102,15 @@ jobs:
type=semver,pattern={{major}}.{{minor}}
type=semver,pattern={{major}}
type=sha
-
name: Set up QEMU
uses: docker/setup-qemu-action@v4
-
name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
uses: docker/setup-buildx-action@f87e5991a6d7451dcb8d9637bfbc97413f497069 # v4
# Pull requests build the image to prove the Dockerfile still works, but
# must not read registry credentials (fork PRs have none).
-
name: Login to Docker Hub
uses: docker/login-action@v4
if: github.event_name != 'pull_request'
uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4
with:
username: ${{ secrets.DOCKERHUB_USERNAME }}
password: ${{ secrets.DOCKERHUB_TOKEN }}
@@ -66,15 +119,26 @@ jobs:
# stage's layers so the cargo-deps and ffmpeg layers are restored
# instead of re-downloaded/recompiled. The scope must be pinned to a
# fixed string: the gha backend defaults to the current git ref, which
# would give every new tag a cold cache on release builds.
# would give every new tag a cold cache on release builds. PR runs only
# read it (cache-to is empty) so they cannot evict the release cache.
#
# FFMPEG_URL/FFMPEG_SHA256 come from repository variables when set, so a
# release can pin an exact ffmpeg build (the Dockerfile default follows
# the project's `/redirect/latest/` URL, which has no sha256 sidecar).
#
# Single-arch (amd64) on purpose: adding arm64 means re-adding
# `docker/setup-qemu-action`, `platforms: linux/amd64,linux/arm64`, and
# parameterizing FFMPEG_URL by $TARGETARCH in the Dockerfile.
-
name: Build and push
uses: docker/build-push-action@v7
uses: docker/build-push-action@c3c9e263c25d99ce0380d002d59b67737d91b0dc # v7
with:
push: true
push: ${{ github.event_name != 'pull_request' }}
build-args: |
APP_NAME=${{ env.APP_NAME }}
FFMPEG_URL=${{ vars.FFMPEG_URL || 'https://ffmpeg.martin-riedl.de/redirect/latest/linux/amd64/release/ffmpeg.zip' }}
FFMPEG_SHA256=${{ vars.FFMPEG_SHA256 }}
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
cache-from: type=gha,scope=tgxmb-build
cache-to: type=gha,mode=max,scope=tgxmb-build
cache-to: ${{ github.event_name != 'pull_request' && 'type=gha,mode=max,scope=tgxmb-build' || '' }}
+29 -21
View File
@@ -2,11 +2,11 @@
## Project Overview
Telegram bot (teloxide) that turns post links from X/Twitter, Pixiv, and Bluesky into media messages (images, video, GIF) with the post's title, author, and tags. It supports batch media splitting, retry with persistence, inline queries, forward-channel rebinding with caption templates, and Pixiv ugoira→MP4 transcoding. README and user-facing strings are in Chinese. The project is a Rust port of a Python predecessor (see `queue.rs` comments referencing `utils/task_queue.py`).
Telegram bot (teloxide) that turns post links from X/Twitter, Pixiv, Bluesky, Misskey (misskey.io), and Bilibili dynamics into media messages (images, video, GIF) with the post's title, author, and tags. It supports batch media splitting, retry with persistence, inline queries, forward-channel rebinding with caption templates, and Pixiv ugoira→MP4 transcoding. README is in Chinese; user-facing bot strings are in English. The project is a Rust port of a Python predecessor (see `queue.rs` comments referencing `utils/task_queue.py`).
Two-crate Cargo workspace (both v1.2.2, edition 2024, resolver 3):
Two-crate Cargo workspace (both v1.7.0, edition 2024, resolver 3):
- **`crates/x-media`** — library that fetches and normalizes media from the three sites. Pure, no Telegram knowledge.
- **`crates/x-media`** — library that fetches and normalizes media from the four sites. Pure, no Telegram knowledge.
- **`crates/xmedia-bot`** — the bot binary: teloxide dispatcher, SQLite-backed chat state, persistent task queue.
## Architecture & Data Flow
@@ -20,21 +20,29 @@ Telegram update → Dispatcher (polling or axum webhook) → dptree branches
Message flow: `message_handler` extracts URLs (from `url`/`text_link` entities, text + caption, deduped) → `x_media::site::fetch(url)``Fetched` → builds a `Task``send::send_media_sequence` (media groups ≤ 9, caption on first item) or `send::send_animation`. On Telegram URL-fetch failure or size error (`send_batch_via_upload`): download via `x_media::site::download_media` to a temp file (≤ 10 MiB), sniff magic bytes (`sniff_ext`), upload via multipart; oversized items fall back to `fallback_url`. On failure: `enqueue_retry` persists resume-state `Task` into the SQLite queue → workers lease (120 s lock TTL) → retry with exponential backoff (≤ 30 s, `MAX_RETRIES = 2`) → dead-letter → `notify_failure`. Success → `post_send_actions`: edit-before-forward prompt with inline buttons, or `copy_messages` to the bound forward channel.
The `x-media` library: `site::fetch(url)` dispatches through the `SITES` registry (per-site `impl Site`, in order twitter → bsky → pixiv) and returns `Ok(None)` for unmatched URLs. `Fetched { source_url, caption, title, media: Vec<Media>, sensitive, site_id, … }`; `caption_with(format)` substitutes `{url} {author} {author_url} {title} {tags}`.
Debug command: `/debug <url>` runs the same `x_media::site::fetch` and replies with `debug_report` (`handlers/commands.rs`) — site id, normalized cache key, source URL, title/author/tags, sensitive flag, caption and the media list — nothing is sent, cached or forwarded; the report is capped at 4000 chars and sent with HTML parse mode: raw fields are escaped, and the caption is wrapped in a `<blockquote>` so it renders exactly like the sent media caption (escaped text and links included).
The `/test <url>` command runs the ordinary link pipeline (`urls::url_media`) with `PostSend::Suppressed`: the media is sent and cached like any other link, but the chat's `forward_channel_id`/`edit_before_forward` are ignored, so a test never forwards to the channel and never opens the edit prompt (retries and dead-letter notifications behave as usual). Both commands use a custom `parse_arg_remainder` parser (whole remainder, trimmed) because teloxide's built-in `split` parser takes exactly one space-separated token.
The `x-media` library: `site::fetch(url)` dispatches through the `SITES` registry (per-site `impl Site`, in order twitter → bsky → misskey → pixiv → bilibili) and returns `Ok(None)` for unmatched URLs. `Fetched { source_url, caption, title, content, media: Vec<Media>, sensitive, site_id, … }` (title and content are split per platform: a pixiv artwork's title and description, a bilibili headline and body, and text-only posts whose text is all `content`); `caption_with(format)` substitutes `{url} {author} {author_url} {title} {content} {tags}`.
## Key Directories
| Path | Purpose |
|---|---|
| `crates/x-media/src/` | Fetch library. `site/mod.rs` = dispatcher + `Fetched`/`FetchError`/`download_media`/`media_size`; `media.rs` = `Media` enum; `examples/fetch.rs` = end-to-end usage sample |
| `crates/x-media/src/site/<twitter\|pixiv\|bsky>/` | One directory per site: `mod.rs` (re-exports), `interface.rs` (PATTERN, `enabled()`, `fetch_from_url()`, `cache_key`/`is_retryable`/`media_headers`, unit struct `<Name>Site` implementing `site::Site`, `From<SiteStruct> for Fetched`), `model.rs` (serde DTOs). Pixiv adds `api.rs` (auth + transport); twitter adds `auth.rs` (logged-in GraphQL `TweetDetail` fallback for NSFW tweets, gated on `TWITTER_AUTH_TOKEN`) |
| `crates/x-media/src/site/<twitter\|pixiv\|bsky\|misskey\|bilibili>/` | One directory per site: `mod.rs` (re-exports), `interface.rs` (PATTERN, `enabled()`, `fetch_from_url()`, `cache_key`/`is_retryable`/`media_headers`, unit struct `<Name>Site` implementing `site::Site`, `From<SiteStruct> for Fetched`), `model.rs` (serde DTOs). Pixiv adds `api.rs` (auth + transport); twitter adds `auth.rs` (logged-in GraphQL `TweetDetail` fallback for NSFW tweets, gated on `TWITTER_AUTH_TOKEN`). Misskey targets misskey.io only (`POST /api/notes/show`, 400+`NO_SUCH_NOTE` → NotFound). Bilibili fetches dynamics (images/animated images only — an attached video degrades to its cover, and its title stands in for the post text, which AV dynamics do not have) from `/x/polymer/web-dynamic/v1/detail` sent with `features=itemOpusStyle` (without that flag the legacy serialization drops an image/text post's body and headline entirely — `desc` comes back `null`; the adapter still parses the legacy `major.draw`/`desc`/`archive` shapes as a fallback). No WBI signature is involved; device cookies `buvid3`/`buvid4` are fetched automatically from `/x/frontend/finger/spi` because bilibili's `-352` risk control starts rejecting plain requests, `BILIBILI_COOKIE` is the escalation when an IP stays blocked; `b23.tv` short links are deliberately unmatched. Twitter's `from_syndication_json` HTML-decodes the API text — syndication and GraphQL `full_text` both arrive pre-escaped (`&gt;` `&lt;` `&amp;` `&#39;`) — so the stored text is raw and the caption escap…
| `crates/xmedia-bot/src/main.rs` | Entry point: env/log init, command registration (`register_commands`), shared `send::BOT` force-init, queue worker start, site login validation (`site::validate_all`), 300 s edit-expiry sweep, dptree handler tree, webhook vs polling dispatch |
| `crates/xmedia-bot/src/config.rs` | Manual env parsing into `Config` |
| `crates/xmedia-bot/src/handlers.rs` | `Command` enum (teloxide `BotCommands`), message/inline/callback handlers, URL extraction, global statics; per-URL work flows through a bounded job channel (256) drained by `URL_WORKERS = 8` workers (`start_url_workers`) — backpressure instead of unbounded spawns (teloxide's per-chat workers are sequential — batch-forwards need concurrency) |
| `crates/xmedia-bot/src/db.rs` | `DbPool`: one shared SQLite connection pool (`POOL_SIZE = 4`, WAL, busy_timeout) for all three tables over `$DATA_DIR/task_queue.db` (default `data/`) — the three stores share it; `open_store` creates file + schema, `with_conn` runs all rusqlite I/O in `spawn_blocking` |
| `crates/xmedia-bot/src/handlers/` | Handler modules: `mod.rs` (message entry point, `reply`, `log_key`), `commands.rs` (teloxide `BotCommands` enum + command executor, incl. `/test <url>` (send-only) / `/debug <url>` (parse-only) and the admin-only `/bot_dict` state dump), `urls.rs` (URL extraction + bounded job channel (256) drained by `URL_WORKERS = 8` workers (`start_url_workers`) — backpressure instead of unbounded spawns; teloxide's per-chat workers are sequential — batch-forwards need concurrency), `inline.rs`/`callback.rs` (inline queries / edit-before-forward buttons), `statics.rs` (global statics) |
| `crates/xmedia-bot/src/state.rs` | `ChatStore`: parking_lot `Mutex<HashMap>` cache + SQLite write-through (`chat_state` table) |
| `crates/xmedia-bot/src/link_cache.rs` | `LinkCache`: SQLite-backed cache (`link_cache` table) of successfully sent posts — raw caption fields + Telegram `file_id`s; repeat links re-send locally (no fetch/upload), TTL + prune, invalidated on permanent send failure |
| `crates/xmedia-bot/src/queue.rs` | `PersistentTaskQueue`: SQLite-backed queue (`tasks` table), `QUEUE_WORKERS = 4` concurrent workers (lease via `BEGIN IMMEDIATE` + `locked_until` TTL), retry→dead-letter, `Notify::notify_waiters` wakeup, `busy_timeout` on all connections |
| `crates/xmedia-bot/src/send.rs` | Media senders, upload fallback, error classification, queue task handlers |
| `crates/xmedia-bot/src/queue.rs` | `PersistentTaskQueue`: SQLite-backed queue (`tasks` table), `QUEUE_WORKERS = 4` concurrent workers (lease via `BEGIN IMMEDIATE` + `locked_until` TTL), retry→dead-letter, `notify_one` worker wakeup plus a separate `Notify` for the 30 s lease-expiry sweep (a shared one let the sweep steal the workers' wakeup permit), `busy_timeout` on all connections |
| `crates/xmedia-bot/src/ctx.rs` | `AppContext`: the injected collaborators (`sender` + `ChatStore`/`PersistentTaskQueue`/`LinkCache`/`Config`), `from_statics` for production and the `CONTEXT` static the worker closures hold. `test_support::TestStores` backs handler tests with a tempdir store set |
| `crates/xmedia-bot/src/send/` | `send/mod.rs`: `Task`/`MediaItemPayload` payloads, `SendError`/`Classification`, `send_media_sequence`/`send_animation`/`forward_messages`; `send/input_media.rs`: payload → `InputFile`/`InputMedia` + `build_media_group` (caption on the first item only); `send/upload.rs`: the download-and-reupload fallback (`prepare_upload_item`/`send_batch_via_upload`, photo downscale handoff); `send/post_send.rs`: link-cache write, `KEEP_ALIVE` registry, `settle_task`, `post_send_actions`, `handle_task`/`dead_letter_notify` |
| `crates/xmedia-bot/src/media_sender.rs` | `MediaSender` trait: the user-flow surface (`send_media_group`/`send_animation`/`copy_messages`/`send_message`/`answer_callback_query`/`edit_message_caption`/`delete_message`/`send_chat_action`) implemented by teloxide `Bot` (per-chat rate-limited) and by a recording `MockSender` in tests. Admin/setup APIs (`get_chat`, `set_my_commands`, …) stay on the concrete `Bot` |
| `crates/xmedia-bot/src/rate_limit.rs` | Per-chat token bucket (`CAPACITY = 20`, ~20 msg/min refill) paced before sends reach the API so batch forwards don't trip flood control |
## Development Commands
@@ -53,13 +61,13 @@ Docker: `docker build -t tgxmb .` then `docker run --rm -d --name tgxmb --env-fi
## Code Conventions & Common Patterns
- **Errors via `thiserror` derive** (no anyhow): the public, stringified errors — `FetchError` (`Http`/`Json`/`Pixiv`/`Site`/`NotFound`/`Blocked`) and `PixivError` — derive `thiserror::Error` with `#[from]` conversions; `Display`/`source()` come from the derive. The internal control-flow enums — `QueueError` (`Retryable { delay_seconds, payload }` / `Permanent`), `SendError` (Retryable/Permanent), `Classification`, `FallbackError` — carry no `Display` and are handled by direct variant matching. New errors should follow the same split: stringified/public errors derive `thiserror`, internal flow enums stay plain.
- **Global state via `std::sync::LazyLock` statics**, not DI: `CONFIG`, `CHAT_STORE`, `TASK_QUEUE` in `handlers.rs`; shared reqwest `CLIENT` in `x-media/src/site/mod.rs`. `Bot` is passed/cloned into handlers; queue workers share the process-wide `send::BOT` (`LazyLock<Bot>`, force-initialized in `main` so a missing token fails at startup).
- **Global state via `std::sync::LazyLock` statics**, not DI: `CONFIG`, `CHAT_STORE`, `TASK_QUEUE` in `handlers/statics.rs`; shared reqwest `CLIENT` in `x-media/src/site/mod.rs`. `Bot` is passed/cloned into handlers; queue workers share the process-wide `send::BOT` (`LazyLock<Bot>`, force-initialized in `main` so a missing token fails at startup).
- **Async**: tokio multi-thread runtime (`#[tokio::main]` default). All rusqlite I/O inside `tokio::task::spawn_blocking`. Long loops use `tokio::select!` with `tokio::sync::{watch, Notify}` stop/wake channels. No streams.
- **Blocking sync primitives**: `parking_lot::Mutex` for hot caches, `tokio::sync::Mutex` for async-shared state (pixiv token cache), `AtomicBool` for feature gates.
- **Site adapter convention**: each site module exports `PATTERN: LazyLock<Regex>`, `enabled() -> bool`, `fetch_from_url(url) -> Result<Fetched, FetchError>`, plus `cache_key`/`is_retryable`/`media_headers`, and a unit struct `<Name>Site` implementing `site::Site`; the central dispatcher (`site/mod.rs`) only iterates the `SITES` registry. Adding a site = new `site/<name>/{mod.rs,interface.rs,model.rs}` + one `Box::new(...)` entry in `SITES` — the bot crate never lists sites (SetFormat whitelist, cache-key site lookup and startup validation all derive from the registry). Async trait methods return `SiteFuture` (a boxed `Pin<Box<dyn Future + Send>>`) because `async fn` in traits is not dyn-compatible.
- **Serde**: per-site `model.rs` are pure `Deserialize` DTOs mirroring API JSON; site structs in `interface.rs` have private fields, a `caption()` builder, and `impl From<SiteStruct> for Fetched`. Persisted payloads use internally-tagged enums (`#[serde(tag = "kind")]` / `type`).
- **Naming**: module-per-concern, snake_case files, `CamelCase` types, `snake_case` fns. `//!` module docs and `///` docs on non-obvious logic (syndication token, ugoira encoding, `display_text_range`).
- **Retries**: only `x-media::site::fetch` retries (3 attempts, `1 << attempt` backoff, HTTP errors only). Queue retries are explicit `QueueError::Retryable` with computed delay (`retry_delay_seconds`).
- **Retries**: only `x-media::site::fetch` retries (3 attempts, `1 << attempt` backoff, HTTP errors only); `site::fetch_once` is the same code path with a single attempt, used by inline queries whose answer window is shorter than the backoff. Queue retries are explicit `QueueError::Retryable` with computed delay (`retry_delay_seconds`).
- Logging via `log` macros (`pretty_env_logger`, level from `RUST_LOG`). Level convention: `info` = lifecycle + per-post business results (`sent`/`forwarded`/`copied`), admin/operator actions and anomalies (fallback, retry enqueue, dead-letter is `error`); `debug` = per-request detail (message/command/URL extraction, `fetching`/`fetched`, batch sends, queue processing, photo processing, inline queries). Full user-submitted URLs and message text only appear at `debug`; at `info` and above links are printed via the normalized cache key (`handlers::log_key`, e.g. `[key=twitter:123...]`) so logs stay short and do not echo user data.
## Important Files
@@ -67,15 +75,15 @@ Docker: `docker build -t tgxmb .` then `docker run --rm -d --name tgxmb --env-fi
| File | Why it matters |
|---|---|
| `crates/xmedia-bot/src/main.rs` | Startup sequence, webhook vs polling, graceful shutdown (SIGINT via teloxide ctrlc / SIGTERM via `stop_token` for docker, → sweep stop → admin msg → queue stop) |
| `crates/xmedia-bot/src/handlers.rs` | `CHAT_STORE`/`TASK_QUEUE`/`CONFIG` singletons (open `data/task_queue.db` **relative to CWD**); command dispatch; URL extraction; retry enqueue |
| `crates/xmedia-bot/src/send.rs` | Constants `MAX_MEDIA_GROUP = 9`; fallback chain; `classify_request_error`; download-and-reupload fallback triggered only by Telegram API errors (`is_media_fetch_failure` / `is_size_error`) |
| `crates/xmedia-bot/src/handlers/` | `statics.rs` = `CHAT_STORE`/`TASK_QUEUE`/`CONFIG` singletons (open `$DATA_DIR/task_queue.db`, default `data/` **relative to CWD**, dir auto-created); `commands.rs` = command dispatch (incl. `/test <url>` send-only, `/debug <url>` parse-only, and the admin-only `/bot_dict` state dump); `urls.rs` = URL extraction + the per-URL pipeline (`url_media` takes a `PostSend` mode: chat settings vs `/test`'s suppressed actions); `inline.rs` = debounced inline queries; `callback.rs` = edit-before-forward buttons (dptree entry + testable `handle_callback` core) |
| `crates/xmedia-bot/src/send/` | `mod.rs`: constants `MAX_MEDIA_GROUP = 9`; `classify_request_error`; the senders. `upload.rs`: download-and-reupload fallback triggered only by Telegram API errors (`is_media_fetch_failure` / `is_size_error`). `post_send.rs`: settlement (`settle_task`), cache write, post-send actions, queue handlers. `input_media.rs`: payload → `InputMedia` |
| `crates/xmedia-bot/src/photo.rs` | Pure-Rust photo processing (no ffmpeg): `png` (image-png) decode/encode + `zune-jpeg` decode + `fast_image_resize` Lanczos3 downscale + `jpeg-encoder`. Photos over Telegram's limits (width + height > 10000 px → `PHOTO_INVALID_DIMENSIONS`; bytes > 10 MiB) are decoded, downscaled keeping the format, PNG bit depth > 24 (RGBA 32-bit / 16-bit per channel) reduced to 24-bit RGB with alpha flattened white (≤24-bit untouched, never upconverted), and transcoded to JPEG only if still over the cap; memory budget guarded, otherwise the item's smaller fallback URL |
| `crates/x-media/src/site/mod.rs` | Dispatcher, `Fetched`/`FetchError`, shared `CLIENT`, `download_media` (adds `Referer: https://www.pixiv.net/` for `pximg.net` hotlink protection) |
| `crates/x-media/src/site/pixiv/api.rs` | OAuth token exchange (hardcoded app client id/secret), access-token cache, ugoira zip→MP4 via ffmpeg in `spawn_blocking` |
| `Dockerfile` | Multi-stage: cached dep layer via stub sources + `touch *.rs` mtime bump (cargo's freshness is mtime-based and `cargo clean -p` removes 0 files — the touch is what forces the real sources to rebuild while deps stay cached), static ffmpeg from ffmpeg.martin-riedl.de (`FFMPEG_URL` arg, optional `FFMPEG_SHA256` checksum, `unzip -t` integrity check), `debian:bookworm-slim` runtime, entrypoint. Runtime ships **no libssl/libcrypto/CA bundle** — rustls webpki-roots handles all TLS, and the static ffmpeg only processes local files (downloads go through reqwest) |
| `docker-entrypoint.sh` | Privilege drop: `useradd` with `LOCAL_USER_ID` (default 9001) + `setpriv` (no gosu on bookworm-slim) |
| `docker-compose.yml.example` | Deployment env reference (real `docker-compose.yml` is gitignored). Ships nginx-proxy + acme-companion: webhook mode needs TLS termination in front (teloxide's axum listener is HTTP-only; `WEBHOOK_CERT` only feeds `set_webhook`), bot exposes `VIRTUAL_HOST`/`VIRTUAL_PORT` on the shared `proxy` network, no host port; container names `nginx-proxy`/`acme-companion`/`tgxmb`, start order via `depends_on` (proxy → acme → bot) |
| `.github/workflows/docker.yml` | CI: build+push to Docker Hub on tag `v*`/master; **no test step**; buildx gha cache (`cache-from`/`cache-to`, scope `tgxmb-build`, `mode=max`) so cargo deps + ffmpeg layers are restored across runs |
| `.github/workflows/docker.yml` | CI: build+push to Docker Hub on tag `v*`/master, plus a build-only check on PRs touching the build inputs; **no test step**; verifies a release tag matches both crate versions; buildx gha cache (`cache-from` always, `cache-to` except on PRs, scope `tgxmb-build`, `mode=max`) so cargo deps + ffmpeg layers are restored across runs; `FFMPEG_URL`/`FFMPEG_SHA256` come from repo variables when set |
| `README.md` | Feature docs + command table (Chinese) |
## Runtime/Tooling Preferences
@@ -83,18 +91,18 @@ Docker: `docker build -t tgxmb .` then `docker run --rm -d --name tgxmb --env-fi
- **Rust, stable, edition 2024**, workspace resolver 3. No `rust-version`/MSRV pin, no `rust-toolchain.toml` — recent stable is assumed. No nightly features.
- Package manager: **Cargo** (workspace with path dep `x-media``xmedia-bot`). No `[workspace.package]`/shared deps — each crate lists deps independently.
- **TLS is rustls end-to-end** (no native-tls/openssl in the tree, no libssl in the Docker runtime image): `teloxide` is declared `default-features = false` with `["webhooks-axum", "macros", "rustls", "ctrlc_handler"]` (the removed `default` also carried `native-tls` and `ctrlc_handler` — the latter must stay); x-media's reqwest is `default-features = false` with `["json", "rustls-tls"]` (webpki-roots baked in, so the image ships no CA bundle). One reqwest 0.12.28 in the lock.
- **Versioning**: bump the version in all three places (`crates/x-media/Cargo.toml`, `crates/xmedia-bot/Cargo.toml`, `Cargo.lock`) and **keep `README.md` and `AGENTS.md` in sync with the code on every bump**, then commit (`chore: bump version to X.Y.Z`), create an annotated tag `vX.Y.Z`, and push branch + tag (the tag push triggers the Docker Hub build).
- Config is **environment-variable driven** (dotenv loads `.env`, gitignored; no `.env.example` exists). Key vars: `TELOXIDE_TOKEN` (required), `PIXIV_REFRESH_TOKEN`, `TWITTER_AUTH_TOKEN` (optional; x.com `auth_token` cookie — enables the logged-in GraphQL fallback that fetches NSFW tweets syndication withholds), `BOT_ADMIN` (comma-separated ids), `EDIT_MESSAGE_TTL_SECONDS` (default 86400), `LINK_CACHE_TTL_SECONDS` (default 604800), `WEBHOOK`/`WEBHOOK_URL`/`WEBHOOK_LISTEN`/`WEBHOOK_PORT`/`WEBHOOK_CERT`/`WEBHOOK_SECRET_TOKEN` (webhook mode requires URL/listen/port, `.expect`ed; `WEBHOOK_CERT` is Telegram-facing self-signed validation only — TLS must be terminated by a reverse proxy), `RUST_LOG`, `TELOXIDE_PROXY`, `LOCAL_USER_ID` (entrypoint only).
- SQLite via `rusqlite` with `bundled` feature (no system libsqlite needed). DB file `data/task_queue.db` is CWD-relative — run from the workspace root, or `/app` in Docker. Mount `./data` and `./cert` volumes.
- **Versioning**: bump the version in all three places (`crates/x-media/Cargo.toml`, `crates/xmedia-bot/Cargo.toml`, `Cargo.lock`) and **keep `README.md`, `README.en.md` and `AGENTS.md` in sync with the code on every bump**, then commit (`chore: bump version to X.Y.Z`), create an annotated tag `vX.Y.Z`, and push branch + tag (the tag push triggers the Docker Hub build). The tag must equal both crate versions: `.github/workflows/docker.yml` verifies that before building, and `--locked` verifies the lock file.
- Config is **environment-variable driven** (dotenv loads `.env`, gitignored; no `.env.example` exists). Key vars: `TELOXIDE_TOKEN` (required), `PIXIV_REFRESH_TOKEN`, `TWITTER_AUTH_TOKEN` (optional; x.com `auth_token` cookie — enables the logged-in GraphQL fallback that fetches NSFW tweets syndication withholds), `BILIBILI_COOKIE` (optional; whole bilibili cookie string — bilibili dynamics fetch anonymously and add their own device cookies, this only rescues an egress IP that bilibili has hard-flagged with `-352`/412), `BOT_ADMIN` (comma-separated ids), `EDIT_MESSAGE_TTL_SECONDS` (default 86400), `LINK_CACHE_TTL_SECONDS` (default 604800), `CAPTION_QUOTE_TEXT_CHARS` (default 200; a post whose text — the `title` plus `content` joined, see `site::compose_text` — reaches this length gets that text wrapped in an expandable blockquote inside its caption, the URL and author line staying outside; `0` disables it. Applied at the send boundary in `send::quote_long_caption`, which locates the text as what follows the author link, so a `/set_format` that moves `{title}`/`{content}` elsewhere and pixiv's title-inside-a-link layout opt out; `copy_messages` forwards and queued retries inherit the wrap, while the edit-before-forward rewrite stays unquoted by design), `DATA_DIR` (default `data`, CWD-relative; the SQLite dir, auto-created), `WEBHOOK`/`WEBHOOK_URL`/`WEBHOOK_LISTEN`/`WEBHOOK_PORT`/`WEBHOOK_CERT`/`WEBHOOK_SECRET_TOKEN` (webhook mode requires URL/listen/port, `.expect`ed; `WEBHOOK_CERT` is Telegram-facing self-signed validation only — TLS must be terminated by a reverse proxy), `RUST_LOG`, `TELOXIDE_PROXY`, `LOCAL_USER_ID` (entrypoint only).
- SQLite via `rusqlite` with `bundled` feature (no system libsqlite needed). DB file `$DATA_DIR/task_queue.db` (default `data/task_queue.db`, CWD-relative — run from the workspace root, or `/app` in Docker; set `DATA_DIR` to pin state anywhere). Mount `./data` and `./cert` volumes.
- `.gitattributes` enforces LF for `*.sh` (CRLF breaks shebangs in containers). `.gitignore`: `.env`, `data/`, `cert/`, `docker-compose.yml`, `/target`, `.idea/`.
- Docs are in Chinese; user-facing bot strings too. Keep that convention when editing captions/templates/docs.
- Docs are in Chinese (README, AGENTS.md); user-facing bot strings are in English. Keep that split when editing user-facing strings and docs.
## Testing & QA
- **~80 tests, all inline `#[cfg(test)] mod tests`** — no `tests/` integration directories. Framework: built-in Rust test + `#[tokio::test]` (dev-deps only in `x-media`: tokio macros/rt-multi-thread, dotenv).
- **~180 tests, all inline `#[cfg(test)] mod tests`** — no `tests/` integration directories. Framework: built-in Rust test + `#[tokio::test]` (dev-deps only in `x-media`: tokio macros/rt-multi-thread, dotenv).
- No mocking framework anywhere (no mockito/wiremock/mockall). Conventions: pure-function units (regex parsing, serde round-trips, chunking, retry math) tested synchronously; async tests use real dependencies — file-backed SQLite via `tempfile` (`queue.rs::new_queue()` helper), live network fetches.
- Live-network tests exist in `site/twitter/interface.rs` (3), `site/bsky/interface.rs` (2), `site/pixiv/interface.rs`/`api.rs`. Test gating convention (enforced by `.github/workflows/ci.yml`): pure unit tests always run; live-network tests carry `#[ignore = "live network: ..."]` (run via `cargo test --workspace -- --ignored live`); token-gated pixiv tests early-return when `PIXIV_REFRESH_TOKEN` is absent **or empty** (an unset GitHub secret arrives as `""``is_err()` alone would run them tokenless and fail). Run the full offline suite with `cargo test --workspace`.
- Live-network tests exist in `site/twitter/interface.rs` (5), `site/bsky/interface.rs` (2), `site/misskey/interface.rs` (1), `site/bilibili/interface.rs` (4), `site/pixiv/api.rs` (1); `photo.rs` adds one `#[ignore = "heavy: …"]` test. `site/mod.rs` also has a **token-gated but not `#[ignore]`d** pixiv download test (`download_media_pixiv_original_with_referer`): it hits `i.pximg.net` whenever `PIXIV_REFRESH_TOKEN` is set, so a local `cargo test --workspace` is not fully offline and can flake on a pixiv CDN body timeout. Test gating convention (enforced by `.github/workflows/ci.yml`): pure unit tests always run; live-network tests carry `#[ignore = "live network: ..."]` (run via `cargo test --workspace -- --ignored live`); token-gated pixiv tests early-return when `PIXIV_REFRESH_TOKEN` is absent **or empty** (an unset GitHub secret arrives as `""``is_err()` alone would run them tokenless and fail), and the bilibili live tests early-return when the API answers risk control (`-352`, which bilibili applies per IP by request volume). Run the full offline suite with `cargo test --workspace`.
- Fixtures are inline `serde_json::json!` builder fns (`fixture()`, `thread_json()`, `illust_json()`), not files. The shared `CLIENT` sets `pool_max_idle_per_host(0)` under `#[cfg(test)]` to avoid cross-runtime `DispatchGone`.
- **CI** — `.github/workflows/ci.yml` runs `cargo fmt --check` + `cargo clippy --workspace --all-targets -- -D warnings` + `cargo test --workspace` (offline, no secrets, on every push/PR) and a `live` job (schedule/manual/tag only, `PIXIV_REFRESH_TOKEN`/`TWITTER_AUTH_TOKEN` from secrets, `continue-on-error`) for the `#[ignore]`d live + token tests. `.github/workflows/docker.yml` builds/pushes the image only.
- Untested and hard to test without a mock seam: `handlers.rs` (depends directly on teloxide `Bot`); `main.rs`, `config.rs`, `state.rs`; `media.rs`, `lib.rs`, all `model.rs`.
- **CI** — `.github/workflows/ci.yml` (actions pinned to commit SHAs, `--locked` on every cargo invocation, `concurrency` cancels superseded runs, `RUST_BACKTRACE=1`) runs `cargo fmt --check` + `cargo clippy --workspace --all-targets --locked -- -D warnings` + `cargo test --workspace --locked` + a release-profile `cargo build --release --locked` + an `actions-rust-lang/audit` dependency-vulnerability gate (offline, no secrets, on every push/PR) and a `live` job (schedule/manual/tag only, `-p x-media` since every network/secret-gated test lives there, `continue-on-error`) for the `#[ignore]`d live + token tests. `.github/workflows/docker.yml` builds and pushes the image on master/tag and runs a **build-only check on pull requests touching the build inputs** (`Dockerfile`, entrypoint, manifests, `.dockerignore`); a release tag must match both crate versions or the build stops, and `FFMPEG_URL`/`FFMPEG_SHA256` are taken from repository variables when set (a release can pin an exact ffmpeg build). `.github/dependabot.yml` keeps crates, the pinned actions and the Docker base images current.
- Untested and hard to test without a mock seam: `main.rs`, `config.rs`, `db.rs`, `handlers/statics.rs`, `media_sender.rs` (holds the `MockSender` itself); in `x-media`: `media.rs`, `lib.rs`, all `model.rs`. The `commands.rs` *executor* needs a real `Bot` (only its pure report builder is tested). Everything else — `handlers/{mod,callback,inline,urls}.rs`, `send/*`, `ctx.rs`, `state.rs`, `queue.rs`, `link_cache.rs`, `rate_limit.rs` — is driven through `TestStores`/`ctx::test_support` and the scripted `MockSender`.
- No coverage tracking.
+168
View File
@@ -0,0 +1,168 @@
# Bilibili 动态支持:研究与实现记录
状态:已实现(`crates/x-media/src/site/bilibili/`)。本文记录上游调研、实测数据与最终设计;
长期契约以 `AGENTS.md` 为准。
范围:**只发动态里的图片与动图**。动态内嵌视频不发流,降级为封面图;`b23.tv` 短链不匹配;
视频页 / 番剧 / 直播间 / 专栏 / 音频均不支持。
---
## 1. 上游实现研究
### 1.1 nazurin`nazurin/sites/bilibili/`4 个文件 ~6 KB
- 入口正则:`t\.bilibili\.com/(\d+)``t\.bilibili\.com/h5/dynamic/detail/(\d+)``bilibili\.com/opus/(\d+)`
- 请求:`GET https://api.bilibili.com/x/polymer/web-dynamic/v1/detail?id={id}`,仅加 `Referer: https://t.bilibili.com/{id}`
**无 cookie、无 WBI 签名、无 `build` 参数**
- 错误:`code == 4101147` → not found`code != 0` 或缺 `data` → 报错。
- 媒体:只取 `item.modules.module_dynamic.major.draw.items[].src`;缩略图 `src + "@518w.jpg"`
`size` 字段单位是 **KB**`major` 为空或 `draw.items` 为空 → "No image found"。
**忽略视频、转发(forward)与纯文字动态**
- caption`"#" + module_author.name` + `module_dynamic.desc.text`,链接写死 `https://www.bilibili.com/opus/{id}`
### 1.2 telegram-bili-feed-helper`biliparser/provider/bilibili/`9 个文件 ~57 KB
- 9 个策略类(Video/Opus/Live/Audio/Read + Feed 基类 + Credential + api 工具):门禁正则
`bilibili\.com|b23\.tv|BV\w{10}|av\d+`,再分流,兜底 `client.head(url)` 跟随重定向后按子串分流。
- 动态:`GET /x/polymer/web-dynamic/desktop/v1/detail?id={id}&build=11605`**单条,无分页**);
客户端带桌面 UA、随机 `buvid3={uuid}infoc`;登录态用 `bilibili-api-python``Credential`
Redis 持久化 `SESSDATA/bili_jct/buvid3/buvid4/ac_time_value/DedeUserID`,扫码登录)。
- **同样没有 WBI 签名 / appkey 签名**playurl 用的是非 WBI 的 `/x/player/playurl`
- 媒体:`major.type` 分派 —— DRAW 取全部 `items[].src`ARCHIVE/PGC/ARTICLE/MUSIC/COMMON/LIVE
只取一张 `cover`;FORWARD 取原动态作者/正文并递归进 `orig` 找媒体。
- 视频:仅独立 video 策略解析(`qn` 720P→480P→360P 试 durl,再退 DASH + ffmpeg 合并);
**动态内嵌视频只发封面**
- 错误:要求 `status==200 && code==0`;风控 `-352`/`-412` 无特殊处理。
### 1.3 取舍
| 维度 | nazurin | bff | 本仓库 |
|---|---|---|---|
| 接口 | `v1/detail?id=` | `desktop/v1/detail?id=&build=` | `v1/detail?id=`(实测可用) |
| 认证 | 无 | buvid3 + SESSDATA | 默认匿名;可选 `BILIBILI_COOKIE` |
| WBI | 无 | 无 | 不实现(无需求) |
| 图片 | `major.draw.items` | 同 + forward 递归 | 同,加 `orig` 递归、`http→https``.gif → Animated` |
| 视频 | 完全忽略 | 动态内嵌视频发封面 | 发封面(不发流) |
| 短链 | 不匹配 | 跟随重定向 | 不匹配(多数短链是视频,会让"静默忽略"变成失败提示) |
---
## 2. 实测验证(2026-09-17,真实请求)
| 验证项 | 结果 |
|---|---|
| `v1/detail?id=`(无 cookie、UA `Mozilla/5.0`、带 Referer | `200 {"code":0}` ✅ |
| 同上,不带 cookie 也不带 Referer | `200 {"code":0}` ✅(无强制鉴权) |
| bff 的 `bilibili_pc/…Electron/22.3.27` UA | `code:-352` ❌ → **不要抄它的 UA** |
| `desktop/v1/detail?build=11605` | `code:-352` ❌ |
| `feed/space?host_mid=`(用户时间线) | 首次成功、随后 `-352`,也见过 HTTP 412 → **不碰** |
| 不存在 / 已删除的动态 | `code:500` "Cannot read property 'only_fans' of undefined"nazurin 的 4101147 已失效) |
| 非数字 id | `code:-400` param parsing failed |
| 图片 `i0.hdslb.com/bfs/new_dyn/*.jpg` | `HEAD 200 image/jpeg`,带/不带 Referer 均可;`+@518w.jpg` → 2542 KB ✅ |
| `t.bilibili.com/h5/dynamic/detail/<id>` | `200` ✅ |
| `m.bilibili.com/dynamic/<id>` | `302 → t.bilibili.com/<id>` ✅ |
| `www.bilibili.com/opus/<id>` | `200`,转发动态 `302 → t.bilibili.com/<id>` ✅ |
| `b23.tv/BV1JTtt6JEZu` | `302 → www.bilibili.com/video/BV…`(视频) |
| `b23.tv/<无效码>` | **HTTP 200** + `{"code":-404}` ⚠️ 短链判定不能只看状态码 |
| `playurl`(仅调研用,未采用) | `fnval=1` 匿名给 durl720P=9.18 MiB / 360P=2.97 MiB`fnval=4048` 匿名 DASH 上限仅 480P |
| `dyn_archive` 字段 | 有 `aid/bvid/cover/title/duration_text`**没有 `cid`**(所以发流要再来一次 `view` 请求) |
| **风控阶梯(同一 IP 连续请求后实测)** | ① 无 cookie → `-352`;② 仅 `buvid3` → 仍 `-352`;③ `buvid3`+`buvid4`(取自匿名 `/x/frontend/finger/spi`)→ **`code:0` 恢复**;④ 继续高频请求后 → 连同 buvid 一起 `-352`(此时只有登录 cookie 或换 IP) |
| **正文位置(24 条真实动态逐条审计)** | 有正文的动态都在 `module_dynamic.desc.text`(图文/转发/纯文字,含 34–193 字样本);**AV(视频投稿)动态 `desc` 恒为 `null`**,内容在 `major.archive.title` / `.desc` 卡片里 → 已做 title 回退 |
| **`features=itemOpusStyle` 的效果** | 同一端点带此参数后,图文帖改为 `major.opus` 形态:`pics[]`(图,key 是 `url`)、`summary.text`(正文,未截断,实测 307 字整段)、`title`(可选标题);不带参数则是 legacy `major.draw` + `desc`,而 **opus 图文帖的 `desc` 为 `null`、正文与标题完全丢失**`opus/1248857553488576532`legacy `desc:null`,带参数 `summary.text="[doge_金箍]黑白搭配"`)。AV / 转发帖不受该参数影响 → 适配器改为请求时带参数,并保留 legacy 形态兜底 |
| feed 与 detail 的差异 | `feed/space` 的 item 会把 `desc.text` 挖空,**只有 detail 有正文** → 排查时不要用 feed 数据判断正文缺失 |
| 不存在的 19 位 id | `4101105 请求数据发生错误`(提示可重试,但只出现在不可能存在的 id 上)→ 仍归入永久错误,见 `code_error` 注释 |
测试样本(live 测试用):
| 样本 | id | 期望 |
|---|---|---|
| 图片动态(2 图 + 话题) | `1245284537985925159` | 2 个 `Illustration``{tags}` = `ALin出道20周年快乐` |
| 转发动态 | `1248982077447077907` | 媒体来自 `orig`1 图),正文可含 `//@` |
| 视频动态 | `1248717597691609105` | 封面 1 张 `Illustration` |
| 纯文字动态 | `1246767523595026450` | `media` 为空 |
关键字段路径:
```
data.item.id_str
data.item.modules.module_author.{name,mid}
data.item.modules.module_dynamic.desc.text
data.item.modules.module_dynamic.topic.{id,name} # 单话题,{tags} 来源
data.item.modules.module_dynamic.major.{draw.items[].src, archive.cover}
data.item.orig # 转发时存在,结构与 item 相同
```
---
## 3. 实现
```
crates/x-media/src/site/bilibili/mod.rs # re-export
crates/x-media/src/site/bilibili/interface.rs # PATTERN / cache_key / enabled / is_retryable /
# media_headers / BilibiliSite / fetch / code_error /
# From<Item> for Fetched / caption / 12 单测 + 2 live
crates/x-media/src/site/bilibili/model.rs # 纯 Deserialize DTO(全 Option
```
- **正则**(同时用于分发、抽 id、缓存键,一个正则三用):
`^(?:https?://)?(?:www|t|m)\.bilibili\.com/(?:opus/|dynamic/|h5/dynamic/detail/)?(\d+)`
- **缓存键**`bilibili:<动态 id>``source_url` 统一 `https://www.bilibili.com/opus/{id}`
- **请求**`GET /x/polymer/web-dynamic/v1/detail?id=` + `Referer: https://www.bilibili.com/`
`Cookie` 头按优先级取:`BILIBILI_COOKIE` → 缓存的设备 cookie`GET /x/frontend/finger/spi``buvid3`/`buvid4`
进程内缓存一次;取不到就不带 cookie,仅 debug 日志)→ 无。指纹接口本身失败**不**让抓取失败。
走共享 `CLIENT`UA `Mozilla/5.0`30s 超时,`TELOXIDE_PROXY` 透传)。
- **错误映射**`0` → 成功;`-352/-412` 与 HTTP 412 → `Transient`(可重试,队列退避;首次记一条 warn 提示
`BILIBILI_COOKIE`);`500`/`4101147``NotFound`(永久);其他 code → `Site`(永久)。
- **媒体**
- `major.opus.pics[]`(带 `features=itemOpusStyle` 时的图文帖形态,字段名是 `url`)→ 每张一张图;
其次 `major.draw.items[]`legacy,字段名 `src`)→ 同样逐张;`http://` / `//``https://`,非 https 开头直接丢弃。
`.gif``Media::Animated``thumbnail_url` 留空,Telegram 自己取首帧——`@518w.jpg` 只对 jpg/webp 实测过),
其余 → `Media::Illustration``thumbnail_url = url + "@518w.jpg"`,兼作超大时的降级 URL)。
- `major.archive.cover` → 1 张 `Illustration`(视频不发流)。
- 转发且自身无媒体 → 递归取 `orig` 的媒体;正文拼 `//@{原作者}:\n{原文}`
- 其他 majorPGC/ARTICLE/MUSIC/LIVE/COMMON)不建模 → 无媒体,走既有 "No media found"。
- **正文 / title**(按信息量从多到少回退):`major.opus.title` + `major.opus.summary.text`
`module_dynamic.desc.text``major.archive.title`。三者分别对应:图文文档(标题+正文)、
legacy/转发帖正文、视频投稿卡片标题。开头结尾空白做 trim;整体再由既有 `truncate_caption` 截断。
- **caption**(与 misskey 同形):`{opus 链接}\n<a href="space.bilibili.com/{mid}">{name}</a>: {正文}`
`RenderData``{tags}` 来自话题名;正文由既有 `truncate_caption` 截断。
- **注册表**`SITES` 末尾追加 → `/set_format` 白名单、链接缓存、启动校验、日志前缀全部自动生效。
- **bot 侧仅文案**`handlers/commands.rs` 三处站点清单字符串 + `state.rs`/`handlers/mod.rs` 注释。
### 与原计划的偏差(及原因)
| 原计划 | 实际 | 原因 |
|---|---|---|
| `x/web-interface/view` + `playurl` 发视频 | 不做 | 需求收窄为图片/动图;视频只发封面 |
| `site/mod.rs``MAX_MEDIA_UPLOAD_BYTES` 常量 | 不加 | 没有视频尺寸决策就不需要该常量,避免跨 crate 耦合 |
| `b23.tv` 短链(跟随重定向) | 不匹配 | 多数短链指向视频,匹配后会把"静默忽略"变成用户的 "Failed to fetch media" |
| `validate()` 校验 cookie | 不做 | 匿名可用,cookie 失效不致命;校验要额外请求一个端点,收益低 |
| `media_headers` 给 hdslb 加 Referer | 返回 `None` | 实测图片与 durl 均无需 Referer(注释里记了这条验证) |
| 计划阶段认为设备 cookie 是 YAGNI,不实现 | **实现**`buvid3`+`buvid4`) | 计划之后做了对照实验:同一 IP 上"无 cookie → -352、只有 buvid3 → -352、buvid3+buvid4 → code:0",说明这是对本适配器主要失败模式的直接修复,而不是冗余保险 |
| 只用不带参数的 `v1/detail` | 加 `features=itemOpusStyle` | 用户实测反馈"有内容的动态没有 title":不带参数时 opus 图文帖返回 legacy 形态,`desc``null`,正文与标题整个丢失。带参数后同一 ID 返回 `major.opus.summary.text` / `title` / `pics`。AV / 转发帖不受影响,legacy 形态仍保留为兜底 |
---
## 4. 测试与验证
- 单元(13):正则匹配/拒绝/忽略短链、缓存键归一、图片映射(https 归一 + 缩略图 + `.gif → Animated`)、
封面、转发取 `orig` 媒体与正文拼接、纯文字无媒体、caption 转义、业务 code 分类(可重试性)、URL 归一、
设备 cookie 拼装。
- live3`#[ignore = "live network: …"]`):设备 cookie 可取、图片动态 2 图、纯文字动态无媒体。
CI 的 `live` job 已覆盖。动态接口被风控时这两条 live 测试打印 `skipping:` 并提前返回(与 pixiv 的
token 门控同款约定),设备 cookie 那条仍会真实执行。
- 实测命令:
`cargo run -p x-media --example fetch -- https://www.bilibili.com/opus/1245284537985925159`
(输出 2 张 `https://i0.hdslb.com/…jpg` + `@518w.jpg` 缩略图 + 话题 tags)。
- 全套:`cargo fmt --check``cargo clippy --workspace --all-targets -- -D warnings``cargo test --workspace` 全绿。
## 5. 已知限制
- 风控按 IP/请求量漂移,阶梯见 §2 最后一行:轻度靠设备 cookie 自愈,重度需 `BILIBILI_COOKIE` 或换 IP。
被拦时按**可重试**失败处理(队列退避)+ 一条 warn,不会静默丢帖。
- 接口 schema 会漂移(`module_dynamic.major` 实测可为 `null` 而正文留在 `desc`);DTO 全 `Option`
未知形态降级为"无媒体",不 panic。
- 动态内嵌视频只发封面图(与 bff 同策略),不下载流。
- 纯文字动态复用既有 "No media found" 回复。
- `b23.tv` 短链不被匹配(见上表)。
Generated
+587 -714
View File
File diff suppressed because it is too large Load Diff
+4 -2
View File
@@ -11,6 +11,8 @@ ARG APP_NAME=telegram-twitter-media-bot
# runners. `/redirect/latest/` floats to the newest release build; each build
# also ships a .sha256. Swap `amd64` for `arm64` when building arm64 images.
ARG FFMPEG_URL=https://ffmpeg.martin-riedl.de/redirect/latest/linux/amd64/release/ffmpeg.zip
# Arm64 images need this URL swapped for the `linux/arm64` build (currently
# hardcoded amd64; the workflow builds amd64 only — see docker.yml).
# Optional sha256 of ffmpeg.zip (pinned releases only): set to verify the
# download. The mirror publishes .sha256 sidecars next to pinned builds, e.g.
# https://ffmpeg.martin-riedl.de/download/linux/amd64/<id>_9.0/ffmpeg.zip.sha256
@@ -28,7 +30,7 @@ COPY crates/xmedia-bot/Cargo.toml crates/xmedia-bot/Cargo.toml
RUN mkdir -p crates/x-media/src crates/xmedia-bot/src \
&& printf 'fn main() {}\n' > crates/xmedia-bot/src/main.rs \
&& : > crates/x-media/src/lib.rs \
&& cargo build --release -p xmedia-bot
&& cargo build --release --locked -p xmedia-bot
# 2. Static ffmpeg next (cached unless FFMPEG_URL changes), so source edits
# never re-download it. The zip contains a single `ffmpeg` binary at the
@@ -51,7 +53,7 @@ RUN wget -q -O /tmp/ffmpeg.zip "$FFMPEG_URL" \
# removes 0 files and the stub binary silently ships.)
COPY crates/ ./crates/
RUN find crates -type f -name '*.rs' -exec touch {} + \
&& cargo build --release -p xmedia-bot
&& cargo build --release --locked -p xmedia-bot
# ---------- runtime stage ----------
FROM debian:bookworm-slim
+12 -4
View File
@@ -1,11 +1,12 @@
# TelegramXMediaBot
A Telegram bot that turns post links from X / Twitter, Pixiv, and Bluesky into media messages (images, video, GIF) with the post's title, author, and tags.
A Telegram bot that turns post links from X / Twitter, Pixiv, Bluesky, Misskey (misskey.io), and Bilibili dynamics into media messages (images, video, GIF) with the post's title, author, and tags.
## Features
- Sending a link in a private chat fetches and sends the images, videos and GIFs automatically; oversized media is split into batches
- Text-only posts report "no media"; unsupported links are silently ignored
- Long posts (text ≥ `CAPTION_QUOTE_TEXT_CHARS`, default 200) show **the text part** of their caption inside a collapsible blockquote, with the link and author line left outside it
- Inline queries (`@bot <link>`)
- Bind a forward channel for automatic forwarding; edit the caption before forwarding and apply custom templates
- Failed sends are retried automatically with persistence; the user is notified after retries are exhausted
@@ -30,10 +31,12 @@ docker build -t tgxmb .
docker run --rm -d --name tgxmb --env-file .env -v ./data:/app/data tgxmb
```
Environment variables: `TELOXIDE_TOKEN` (required), `PIXIV_REFRESH_TOKEN`, `BOT_ADMIN`, `EDIT_MESSAGE_TTL_SECONDS`, `LINK_CACHE_TTL_SECONDS`, `RUST_LOG`, `WEBHOOK*`, `TWITTER_AUTH_TOKEN` (optional).
Environment variables: `TELOXIDE_TOKEN` (required), `PIXIV_REFRESH_TOKEN`, `BOT_ADMIN`, `EDIT_MESSAGE_TTL_SECONDS`, `LINK_CACHE_TTL_SECONDS`, `RUST_LOG`, `TELOXIDE_PROXY`, `WEBHOOK*`, `TWITTER_AUTH_TOKEN` (optional), `BILIBILI_COOKIE` (optional).
NSFW tweets: the public syndication endpoint does not return sensitive content. Setting `TWITTER_AUTH_TOKEN` (the `auth_token` cookie value of a logged-in x.com session) lets the bot fetch NSFW media in the logged-in state only when it hits a withheld tweet; without it, the bot reports no media.
Bilibili dynamics are fetched anonymously by default (no login; the bot fetches bilibili's anonymous `buvid3`/`buvid4` device cookies itself to raise the success rate). If the server's egress IP gets hard-flagged by bilibili (persistent `risk control (-352)` log lines or HTTP 412), set `BILIBILI_COOKIE` (the whole cookie string from a logged-in browser, e.g. `SESSDATA=…; bili_jct=…`) to restore access. Only a dynamic's images and animations are sent; an attached video degrades to its cover image.
### Webhook deployment (needs a reverse proxy)
`docker-compose.yml.example` ships an [nginx-proxy](https://github.com/nginx-proxy/nginx-proxy) + [acme-companion](https://github.com/nginx-proxy/acme-companion) reverse-proxy orchestration. Pick one deployment shape:
@@ -81,9 +84,12 @@ Telegram only accepts ports 443/80/88/8443.
|---|---|
| `TELOXIDE_TOKEN` | Bot token (required) |
| `PIXIV_REFRESH_TOKEN` | Pixiv refresh token; Pixiv is disabled without it |
| `BILIBILI_COOKIE` | Optional bilibili cookie string (`SESSDATA=…; bili_jct=…`); only needed when the egress IP stays risk-controlled (device cookies are fetched automatically) |
| `BOT_ADMIN` | Admin chat IDs, comma-separated; receives start/stop notifications |
| `EDIT_MESSAGE_TTL_SECONDS` | Edit-before-forward record expiry in seconds, default 86400 |
| `LINK_CACHE_TTL_SECONDS` | Link-result cache expiry in seconds, default 604800 (7 days) |
| `CAPTION_QUOTE_TEXT_CHARS` | **The text part** of the caption (the joined `{title}` + `{content}`) is wrapped in a collapsible blockquote once it reaches this many characters, default 200; `0` disables |
| `DATA_DIR` | Data directory (where the SQLite `task_queue.db` lives), default `data` (relative to the working directory, created automatically) |
| `RUST_LOG` | Log level |
| `TELOXIDE_PROXY` | HTTP proxy (e.g. `http://127.0.0.1:10808`); applies to both the Telegram Bot API and site fetches — required on restricted networks (e.g. behind the GFW) |
| `LOCAL_USER_ID` | UID the container runs as, default 9001 |
@@ -110,9 +116,11 @@ Telegram only accepts ports 443/80/88/8443.
| `/remove_forward_channel` | Remove the forward channel |
| `/edit_before_forward` | Toggle "edit before forward": when enabled, the bot posts a prompt after forwarding; replying to it edits the first forwarded message's caption (or taps a template button to apply one) |
| `/set_template <name>` | Reply to a message containing `[]` to save it as a named template; `[]` is replaced by the original post link when forwarding (used with "edit before forward") |
| `/set_format <site> <format>` | Customize the caption format for one site. Sites: `twitter` / `bsky` / `pixiv`. Placeholders: `{url}` `{author}` `{author_url}` `{title}` `{tags}` |
| `/set_format <site> <format>` | Customize the caption format for one site. Sites: `twitter` / `bsky` / `pixiv` / `misskey` / `bilibili`. Placeholders: `{url}` `{author}` `{author_url}` `{title}` `{content}` `{tags}` |
| `/clear_cache [link]` | Clear the link cache (admin only); with a link only that entry, otherwise everything |
| `/bot_dict` | Show the current chat state (debugging) |
| `/bot_dict` | Show the current chat state (debugging; admin only) |
| `/test <link>` | Parse a link and send its media; no channel forward, no edit-before-forward prompt (send only) |
| `/debug <link>` | Debug: parse a link and report the parse result only (site, title, author, tags, media list) — no media is sent |
Link processing works only in private chats; commands work in any chat.
+12 -4
View File
@@ -1,11 +1,12 @@
# TelegramXMediaBot
Telegram 机器人,将 X / Twitter、Pixiv、Bluesky 的帖子链接转换为媒体消息发送,附带帖子标题、作者与标签。
Telegram 机器人,将 X / Twitter、Pixiv、Bluesky、Misskey (misskey.io)、Bilibili 动态的帖子链接转换为媒体消息发送,附带帖子标题、作者与标签。
## 功能
- 私聊发送链接后自动抓取并发送图片、视频与 GIF,超量图片自动分批
- 纯文字帖提示无媒体;不支持的链接静默忽略
- 长帖(正文 ≥ `CAPTION_QUOTE_TEXT_CHARS`,默认 200)的**正文部分**用可折叠引用块展示,链接与作者行留在引用块外
- 支持内联查询(`@机器人 <链接>`
- 可绑定转发频道自动转发;支持转发前编辑 caption 与自定义模板
- 发送失败自动重试并持久化,重试耗尽后通知用户
@@ -30,10 +31,12 @@ docker build -t tgxmb .
docker run --rm -d --name tgxmb --env-file .env -v ./data:/app/data tgxmb
```
环境变量:`TELOXIDE_TOKEN`(必填)、`PIXIV_REFRESH_TOKEN``BOT_ADMIN``EDIT_MESSAGE_TTL_SECONDS``LINK_CACHE_TTL_SECONDS``RUST_LOG``WEBHOOK*``TWITTER_AUTH_TOKEN`(可选)。
环境变量:`TELOXIDE_TOKEN`(必填)、`PIXIV_REFRESH_TOKEN``BOT_ADMIN``EDIT_MESSAGE_TTL_SECONDS``LINK_CACHE_TTL_SECONDS``RUST_LOG``TELOXIDE_PROXY``WEBHOOK*``TWITTER_AUTH_TOKEN`(可选)`BILIBILI_COOKIE`(可选)
NSFW 推文:公开的 syndication 接口不返回敏感内容。设置 `TWITTER_AUTH_TOKEN`(登录 x.com 后浏览器 Cookie 里的 `auth_token` 值)后,bot 会仅在遇到 NSFW 推文时以登录态获取媒体;未设置则提示无媒体。
Bilibili 动态默认匿名抓取(无需登录,bot 会自动从 B 站的匿名指纹接口取 `buvid3`/`buvid4` 设备 cookie 以提高成功率)。若服务器出口 IP 被 B 站重度风控(日志里的 `risk control (-352)` 或 HTTP 412,且持续出现),设置 `BILIBILI_COOKIE`(登录后浏览器里整条 Cookie 串,如 `SESSDATA=…; bili_jct=…`)可恢复访问。当前只发送动态里的图片与动图,动态内嵌视频发送其封面。
### Webhook 部署(需要反向代理)
`docker-compose.yml.example` 内置了 [nginx-proxy](https://github.com/nginx-proxy/nginx-proxy) + [acme-companion](https://github.com/nginx-proxy/acme-companion) 反向代理编排,按部署环境二选一:
@@ -81,9 +84,12 @@ Telegram 只接受 443/80/88/8443 端口。
|---|---|
| `TELOXIDE_TOKEN` | Bot token(必填) |
| `PIXIV_REFRESH_TOKEN` | Pixiv 刷新令牌;未设置则禁用 Pixiv |
| `BILIBILI_COOKIE` | 可选的 B 站 Cookie 串(`SESSDATA=…; bili_jct=…`),仅在出口 IP 被持续风控时才需要(设备 cookie 由 bot 自动获取) |
| `BOT_ADMIN` | 管理员聊天 ID,逗号分隔;接收启动/停止通知 |
| `EDIT_MESSAGE_TTL_SECONDS` | 转发前编辑记录过期秒数,默认 86400 |
| `LINK_CACHE_TTL_SECONDS` | 链接结果缓存过期秒数,默认 604800(7 天) |
| `CAPTION_QUOTE_TEXT_CHARS` | 正文(`{title}` + `{content}` 合计)达到该长度(字符)时,caption 的**正文部分**用可折叠引用块包裹,默认 200;`0` 关闭 |
| `DATA_DIR` | 数据目录(SQLite 数据库 `task_queue.db` 所在目录),默认 `data`(相对工作目录,会自动创建) |
| `RUST_LOG` | 日志级别 |
| `TELOXIDE_PROXY` | HTTP 代理(如 `http://127.0.0.1:10808`);同时作用于 Telegram Bot API 与站点抓取请求,网络受限环境(如 GFW)必需 |
| `LOCAL_USER_ID` | 容器内运行用户 UID,默认 9001 |
@@ -110,9 +116,11 @@ Telegram 只接受 443/80/88/8443 端口。
| `/remove_forward_channel` | 取消转发频道 |
| `/edit_before_forward` | 开关「转发前编辑」:开启后,转发成功后 bot 会发一条提示消息,回复它可修改第一条转发消息的 caption(或点击模板按钮套用模板) |
| `/set_template <名称>` | 回复一条含 `[]` 的消息,将其保存为命名模板;转发时 `[]` 会被替换为原帖链接(配合「转发前编辑」使用) |
| `/set_format <站点> <格式>` | 自定义某站点的 caption 格式。站点:`twitter` / `bsky` / `pixiv`。占位符:`{url}` `{author}` `{author_url}` `{title}` `{tags}` |
| `/set_format <站点> <格式>` | 自定义某站点的 caption 格式。站点:`twitter` / `bsky` / `pixiv` / `misskey` / `bilibili`。占位符:`{url}` `{author}` `{author_url}` `{title}` `{content}` `{tags}` |
| `/clear_cache [链接]` | 清空链接缓存(仅管理员);带链接只清该条,否则清空全部 |
| `/bot_dict` | 查看当前聊天状态(调试用) |
| `/bot_dict` | 查看当前聊天状态(调试用;仅管理员 |
| `/test <链接>` | 解析链接并发送媒体;不转发到频道、不弹转发前编辑提示(仅发送) |
| `/debug <链接>` | 调试:只解析链接并返回解析结果(站点、标题、作者、标签、媒体列表),不发送任何媒体 |
链接处理仅限私聊;命令在任意聊天可用。
+3 -3
View File
@@ -1,6 +1,6 @@
[package]
name = "x-media"
version = "1.2.2"
version = "1.7.0"
edition = "2024"
[dependencies]
@@ -11,10 +11,10 @@ regex = "1.12"
html-escape = "0.2"
url = "2.5.2"
bytes = "1"
zip = "2"
zip = "8"
tempfile = "3"
thiserror = "2"
rand = "0.8"
rand = "0.10"
log = "0.4"
tokio = { version = "1.40", features = ["time"] }
File diff suppressed because it is too large Load Diff
+6
View File
@@ -0,0 +1,6 @@
mod interface;
mod model;
pub use interface::{
BilibiliSite, PATTERN, cache_key, enabled, fetch_from_url, is_retryable, media_headers,
};
+158
View File
@@ -0,0 +1,158 @@
//! Serde DTOs for the Bilibili dynamic detail endpoint
//! (`/x/polymer/web-dynamic/v1/detail`), mirroring live responses
//! (field paths verified 2026-09-17). Every field is optional so an API
//! shape change degrades to "no media" instead of a parse failure.
use serde::Deserialize;
#[derive(Deserialize, Debug)]
pub(crate) struct Detail {
/// Business code: `0` = OK, `-352`/`-412` = risk control, `500`/`4101147`
/// = gone.
pub(crate) code: i64,
#[serde(default)]
pub(crate) message: Option<String>,
#[serde(default)]
pub(crate) data: Option<Data>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Data {
#[serde(default)]
pub(crate) item: Option<Box<Item>>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Item {
/// The dynamic id, same numeric id as in the URL.
#[serde(default)]
pub(crate) id_str: String,
#[serde(default)]
pub(crate) modules: Option<Modules>,
/// The quoted dynamic when this item is a forward. A forward shell often
/// carries no media of its own — the original holds it.
#[serde(default)]
pub(crate) orig: Option<Box<Item>>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Modules {
#[serde(default)]
pub(crate) module_author: Option<Author>,
#[serde(default)]
pub(crate) module_dynamic: Option<Dynamic>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Author {
#[serde(default)]
pub(crate) name: String,
#[serde(default)]
pub(crate) mid: Option<i64>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Dynamic {
#[serde(default)]
pub(crate) desc: Option<Desc>,
#[serde(default)]
pub(crate) major: Option<Major>,
/// A single topic (`{"id":…,"name":…}`), the dynamic's only tag source.
#[serde(default)]
pub(crate) topic: Option<Topic>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Desc {
#[serde(default)]
pub(crate) text: String,
}
/// `major` is a tagged union: `type` (`MAJOR_TYPE_DRAW` / `_OPUS` /
/// `_ARCHIVE` / …) plus one payload object per type. Only the three payloads
/// this adapter reads are modeled; an unknown major simply yields no media.
#[derive(Deserialize, Debug)]
pub(crate) struct Major {
#[serde(default)]
pub(crate) draw: Option<Draw>,
#[serde(default)]
pub(crate) opus: Option<Opus>,
#[serde(default)]
pub(crate) archive: Option<Archive>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Draw {
#[serde(default)]
pub(crate) items: Vec<Pic>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Pic {
/// `major.draw` image URL.
#[serde(default)]
pub(crate) src: Option<String>,
/// `major.opus.pics` image URL — the opus shape names the field
/// differently while carrying the same image.
#[serde(default)]
pub(crate) url: Option<String>,
}
impl Pic {
/// The image URL, whichever key this serialization put it under.
pub(crate) fn url(&self) -> Option<&str> {
self.src.as_deref().or(self.url.as_deref())
}
}
/// `major.opus`: the serialization of an image/text post the web client asks
/// for (`features=itemOpusStyle`). It carries the parts the legacy shape drops
/// entirely — the document title and body of an opus post, whose
/// `module_dynamic.desc` comes back `null`.
#[derive(Deserialize, Debug)]
pub(crate) struct Opus {
/// Document headline; often absent.
#[serde(default)]
pub(crate) title: Option<String>,
/// Document body (untruncated: a 307-char sample came back whole).
#[serde(default)]
pub(crate) summary: Option<Desc>,
#[serde(default)]
pub(crate) pics: Vec<Pic>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Archive {
/// The attached video's cover — the only image an AV dynamic has (the
/// video itself is deliberately not resolved, see the module docs).
#[serde(default)]
pub(crate) cover: Option<String>,
/// The video's title. An AV dynamic has no body of its own (`desc` comes
/// back `null`), so this card title is the post's content.
#[serde(default)]
pub(crate) title: Option<String>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct Topic {
#[serde(default)]
pub(crate) name: String,
}
/// Response of the anonymous fingerprint endpoint (`/x/frontend/finger/spi`),
/// the source of the adapter's device cookies.
#[derive(Deserialize, Debug)]
pub(crate) struct Fingerprint {
#[serde(default)]
pub(crate) data: Option<FingerprintData>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct FingerprintData {
/// Sent as the `buvid3` cookie.
#[serde(default, rename = "b_3")]
pub(crate) buvid3: String,
/// Sent as the `buvid4` cookie.
#[serde(default, rename = "b_4")]
pub(crate) buvid4: String,
}
+7 -3
View File
@@ -330,13 +330,16 @@ impl From<Post> for Fetched {
url: url.clone(),
author: encode_text(&post.author).into_owned(),
author_url: author_url.clone(),
title: encode_text(&post.text).into_owned(),
// A post has no title: its text is all content.
title: String::new(),
content: encode_text(&post.text).into_owned(),
tags: String::new(),
});
Fetched {
source_url: url,
caption: post.caption(),
title: post.text.clone(),
title: String::new(),
content: post.text.clone(),
media: post.media,
sensitive: post.sensitive,
site_id: "bsky",
@@ -410,7 +413,8 @@ mod tests {
fetched.source_url,
"https://bsky.app/profile/user.bsky.social/post/3xxxx"
);
assert_eq!(fetched.title, "hello <world>");
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "hello <world>");
assert_eq!(fetched.media.len(), 1);
assert!(!fetched.sensitive);
// display_name absent -> empty fallback
@@ -0,0 +1,390 @@
//! Site adapter for misskey.io notes: URL pattern, API fetch and
//! normalization into [`Fetched`] (see [`crate::site::Site`]).
use super::model;
use crate::media::Media;
use crate::site::{FetchError, Fetched, RenderData, Site, SiteFuture};
use html_escape::{encode_double_quoted_attribute, encode_text};
use regex::Regex;
use std::sync::LazyLock;
const API_URL: &str = "https://misskey.io/api/notes/show";
/// Registry entry for the misskey.io adapter (see [`crate::site::Site`]).
pub struct MisskeySite;
impl Site for MisskeySite {
fn id(&self) -> &'static str {
"misskey"
}
fn pattern(&self) -> &'static Regex {
&PATTERN
}
fn cache_key(&self, url: &str) -> Option<String> {
cache_key(url)
}
fn fetch_from_url<'a>(&'a self, url: &'a str) -> SiteFuture<'a, Fetched> {
Box::pin(async move { fetch_from_url(url).await })
}
}
pub static PATTERN: LazyLock<Regex> =
LazyLock::new(|| Regex::new(r"^(?:https?://)?misskey\.io/notes/([\w.\-~]+)").unwrap());
pub fn enabled() -> bool {
true
}
pub async fn fetch_from_url(url: &str) -> Result<Fetched, FetchError> {
let caps = PATTERN.captures(url).ok_or(FetchError::NotFound)?;
let note_id = caps.get(1).ok_or(FetchError::NotFound)?.as_str();
let note = fetch(note_id).await?;
Ok(note.into())
}
/// Cache key for a misskey URL: `"misskey:<note id>"`. The prefix is the
/// site id used for caption-format lookup and link-cache keys.
pub fn cache_key(url: &str) -> Option<String> {
PATTERN
.captures(url)
.map(|caps| format!("misskey:{}", &caps[1]))
}
/// Misskey's fetch-retry policy: transient classes only. Not-found, blocked
/// and parse failures are permanent.
pub fn is_retryable(err: &FetchError) -> bool {
matches!(err, FetchError::Http(_) | FetchError::Transient(_))
}
/// misskey.io media hosts need no extra headers (verified: direct GET works).
pub fn media_headers(_url: &str) -> Option<Vec<(&'static str, String)>> {
None
}
/// Fetches a note from misskey.io by id. The API answers client failures
/// with HTTP 400 + `{"error":{"code":...}}` (NO_SUCH_NOTE → NotFound);
/// everything else non-success is transient and retried by [`crate::site::fetch`].
pub async fn fetch(note_id: &str) -> Result<model::Note, FetchError> {
let response = crate::site::CLIENT
.post(API_URL)
.json(&serde_json::json!({ "noteId": note_id }))
.send()
.await?;
let status = response.status();
if !status.is_success() {
return Err(match status.as_u16() {
400 => not_found_or_invalid(response).await,
_ => FetchError::Transient(format!("misskey status {status}")),
});
}
response.json().await.map_err(|e| FetchError::Site {
site: "misskey",
error: Box::new(e),
})
}
/// Maps a 400 response: NO_SUCH_NOTE is permanent NotFound, any other 400 is
/// a site error (permanent — retrying a rejected request cannot succeed).
async fn not_found_or_invalid(response: reqwest::Response) -> FetchError {
match response.json::<serde_json::Value>().await {
Ok(v) if v["error"]["code"] == "NO_SUCH_NOTE" => FetchError::NotFound,
_ => FetchError::Site {
site: "misskey",
error: "note rejected (invalid param or private note)".into(),
},
}
}
/// The note whose content matters: a renote shell has no text/files of its
/// own — the embedded renote carries them.
fn effective(note: &model::Note) -> &model::Note {
match &note.renote {
Some(renote) if note.files.is_empty() => renote,
_ => note,
}
}
impl From<model::Note> for Fetched {
fn from(note: model::Note) -> Self {
let note = &note;
let content = effective(note);
let url = format!("https://misskey.io/notes/{}", note.id);
let author = content
.user
.name
.as_deref()
.filter(|n| !n.is_empty())
.unwrap_or(&content.user.username)
.to_string();
let author_url = format!("https://misskey.io/@{}", content.user.username);
let cw = content.cw.as_deref().unwrap_or_default();
// Notes carry hashtags inline in the text (no structured tags array);
// a CW note gets the marker prefixed so recipients see the spoiler.
let mut text = cw.to_string();
if !cw.is_empty() && !text.ends_with(' ') {
text.push(' ');
}
text.push_str(content.text.as_deref().unwrap_or_default().trim());
let text = text.trim().to_string();
let caption = caption(&url, &author_url, &author, &text);
let sensitive = content.cw.is_some() || content.files.iter().any(|f| f.is_sensitive);
let media: Vec<Media> = content.files.iter().filter_map(media_from_file).collect();
Fetched {
source_url: url.clone(),
caption,
// A note has no title: its text (CW marker included) is content.
title: String::new(),
content: text.clone(),
media,
sensitive,
site_id: "misskey",
render_data: Some(RenderData {
url,
author: encode_text(&author).into_owned(),
author_url: author_url.clone(),
title: String::new(),
content: encode_text(&text).into_owned(),
tags: String::new(),
}),
_keep_alive: None,
}
}
}
fn caption(url: &str, author_url: &str, author: &str, text: &str) -> String {
let url = encode_double_quoted_attribute(url);
let author_url = encode_double_quoted_attribute(author_url);
let author = encode_text(author);
if text.is_empty() {
return format!("{url}\n<a href=\"{author_url}\">{author}</a>");
}
format!(
"{url}\n<a href=\"{author_url}\">{author}</a>: {text}",
text = encode_text(text),
)
}
/// Maps a Misskey DriveFile to a [`Media`] item; unknown/audio/other types
/// are skipped (twitter's `_ => {}` precedent). GIF must be matched before
/// the generic image arm.
fn media_from_file(file: &model::DriveFile) -> Option<Media> {
let title = file.name.clone();
match file.mime_type.as_str() {
"image/gif" => Some(Media::Animated {
title,
url: file.url.clone(),
thumbnail_url: file.thumbnail_url.clone().unwrap_or_default(),
}),
mime if mime.starts_with("image/") => Some(Media::Illustration {
title,
url: file.url.clone(),
thumbnail_url: file.thumbnail_url.clone(),
fallback_url: None,
}),
mime if mime.starts_with("video/") => Some(Media::Video {
title,
url: file.url.clone(),
thumbnail_url: file.thumbnail_url.clone().unwrap_or_default(),
}),
_ => None,
}
}
#[cfg(test)]
mod tests {
use super::*;
fn note_json(json: serde_json::Value) -> model::Note {
serde_json::from_value(json).unwrap()
}
fn base_note() -> serde_json::Value {
serde_json::json!({
"id": "aotihl10lqrs015s",
"text": "hello",
"user": { "name": "ミロン", "username": "donyan47897", "host": null },
"files": []
})
}
#[test]
fn pattern_matches_misskey_note_urls() {
for url in [
"https://misskey.io/notes/aotihl10lqrs015s",
"http://misskey.io/notes/aotihl10lqrs015s",
"misskey.io/notes/aotihl10lqrs015s",
] {
assert!(PATTERN.is_match(url), "{url}");
}
for url in [
"https://misskey.io/",
"https://misskey.io/@user",
"https://misskey.io/notes/",
"https://x.com/user/status/123",
] {
assert!(!PATTERN.is_match(url), "{url}");
}
}
#[test]
fn cache_key_normalizes_variants() {
assert_eq!(
cache_key("https://misskey.io/notes/aotihl10lqrs015s"),
Some("misskey:aotihl10lqrs015s".to_string())
);
assert_eq!(x_media_site_id("misskey:abc"), "misskey");
}
fn x_media_site_id(key: &str) -> &'static str {
crate::site::site_id_from_key(key)
}
#[test]
fn from_json_image_file() {
let mut note = base_note();
note["files"] = serde_json::json!([{
"type": "image/webp",
"url": "https://media.misskeyusercontent.jp/io/a.webp",
"thumbnailUrl": "https://media.misskeyusercontent.jp/io/t.webp",
"isSensitive": true,
"name": "pic.webp"
}]);
let fetched: Fetched = note_json(note).into();
assert_eq!(
fetched.source_url,
"https://misskey.io/notes/aotihl10lqrs015s"
);
assert_eq!(fetched.site_id, "misskey");
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "hello");
assert!(fetched.sensitive);
assert_eq!(fetched.media.len(), 1);
match &fetched.media[0] {
Media::Illustration {
title,
url,
thumbnail_url,
fallback_url,
} => {
assert_eq!(title.as_deref(), Some("pic.webp"));
assert_eq!(url, "https://media.misskeyusercontent.jp/io/a.webp");
assert_eq!(
thumbnail_url.as_deref(),
Some("https://media.misskeyusercontent.jp/io/t.webp")
);
assert!(fallback_url.is_none());
}
other => panic!("expected illustration, got {other:?}"),
}
}
#[test]
fn from_json_gif_video_and_skip_audio() {
let mut note = base_note();
note["files"] = serde_json::json!([
{ "type": "audio/mpeg", "url": "https://m/a.mp3", "isSensitive": false },
{ "type": "image/gif", "url": "https://m/a.gif", "isSensitive": false },
{ "type": "video/webm", "url": "https://m/a.webm", "isSensitive": false }
]);
let fetched: Fetched = note_json(note).into();
assert_eq!(fetched.media.len(), 2);
assert!(
matches!(&fetched.media[0], Media::Animated { url, .. } if url == "https://m/a.gif")
);
assert!(matches!(&fetched.media[1], Media::Video { url, .. } if url == "https://m/a.webm"));
// No thumbnailUrl → empty string, not a broken URL.
match &fetched.media[1] {
Media::Video { thumbnail_url, .. } => assert_eq!(thumbnail_url, ""),
other => panic!("expected video, got {other:?}"),
}
assert!(!fetched.sensitive);
}
#[test]
fn from_json_cw_marks_sensitive_and_prefixes_title() {
let mut note = base_note();
note["cw"] = serde_json::json!("spoiler");
note["text"] = serde_json::json!("body");
let fetched: Fetched = note_json(note).into();
assert!(fetched.sensitive);
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "spoiler body");
}
#[test]
fn from_json_author_falls_back_to_username() {
let mut note = base_note();
note["user"] = serde_json::json!({ "name": null, "username": "donyan47897", "host": null });
let fetched: Fetched = note_json(note).into();
assert!(
fetched.caption.contains("donyan47897"),
"{}",
fetched.caption
);
assert!(fetched.caption.contains("https://misskey.io/@donyan47897"));
}
#[test]
fn from_json_renote_uses_embedded_content() {
let note = serde_json::json!({
"id": "shell0000000000",
"text": null,
"user": { "name": "shell", "username": "shelluser", "host": null },
"files": [],
"renote": {
"id": "inner000000000",
"text": "inner text",
"user": { "name": "inner", "username": "inneruser", "host": null },
"files": [
{ "type": "image/png", "url": "https://m/i.png", "isSensitive": false }
]
}
});
let fetched: Fetched = note_json(note).into();
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "inner text");
assert_eq!(fetched.media.len(), 1);
// The source URL still points at the renote shell the user posted.
assert_eq!(
fetched.source_url,
"https://misskey.io/notes/shell0000000000"
);
}
#[test]
fn caption_layout_matches_bsky() {
let fetched: Fetched = note_json(base_note()).into();
assert_eq!(
fetched.caption,
"https://misskey.io/notes/aotihl10lqrs015s\n<a href=\"https://misskey.io/@donyan47897\">ミロン</a>: hello"
);
}
#[test]
fn caption_without_text_has_no_dangling_colon() {
let mut note = base_note();
note["text"] = serde_json::json!(null);
let fetched: Fetched = note_json(note).into();
assert_eq!(
fetched.caption,
"https://misskey.io/notes/aotihl10lqrs015s\n<a href=\"https://misskey.io/@donyan47897\">ミロン</a>"
);
}
#[tokio::test]
#[ignore = "live network: requires outbound HTTPS to misskey.io"]
async fn live_fetch_reference_note() {
let fetched = fetch_from_url("https://misskey.io/notes/aotihl10lqrs015s")
.await
.unwrap();
assert_eq!(fetched.site_id, "misskey");
assert_eq!(fetched.media.len(), 1);
assert!(fetched.sensitive);
assert!(!fetched.caption.is_empty());
}
}
+6
View File
@@ -0,0 +1,6 @@
mod interface;
mod model;
pub use interface::{
MisskeySite, PATTERN, cache_key, enabled, fetch_from_url, is_retryable, media_headers,
};
+35
View File
@@ -0,0 +1,35 @@
use serde::Deserialize;
#[derive(Deserialize, Debug)]
pub(crate) struct Note {
pub(crate) id: String,
pub(crate) text: Option<String>,
#[serde(default)]
pub(crate) cw: Option<String>,
pub(crate) user: User,
#[serde(default)]
pub(crate) files: Vec<DriveFile>,
/// Embedded original note when this note is a renote; the shell's own
/// text/files are usually empty and the content lives here.
#[serde(default)]
pub(crate) renote: Option<Box<Note>>,
}
#[derive(Deserialize, Debug)]
pub(crate) struct User {
pub(crate) name: Option<String>,
pub(crate) username: String,
}
#[derive(Deserialize, Debug)]
pub(crate) struct DriveFile {
#[serde(rename = "type")]
pub(crate) mime_type: String,
pub(crate) url: String,
#[serde(default, rename = "thumbnailUrl")]
pub(crate) thumbnail_url: Option<String>,
#[serde(default, rename = "isSensitive")]
pub(crate) is_sensitive: bool,
#[serde(default)]
pub(crate) name: Option<String>,
}
+107 -31
View File
@@ -1,8 +1,8 @@
//! Site fetching dispatcher and unified result types.
//!
//! Dispatch order: twitter → bsky → pixiv. Each site module exports a
//! `PATTERN`, `enabled()` and `fetch_from_url()`; a future site plugs in by
//! adding one guarded entry in [`fetch_once`].
//! Dispatch order: twitter → bsky → misskey → pixiv → bilibili. Each site
//! module exports a `PATTERN`, `enabled()` and `fetch_from_url()`; a future
//! site plugs in by adding one guarded entry in `SITES`.
use std::future::Future;
use std::pin::Pin;
@@ -13,28 +13,39 @@ use std::time::Duration;
use regex::Regex;
use thiserror::Error;
pub mod bilibili;
pub mod bsky;
pub mod misskey;
pub mod pixiv;
pub mod twitter;
pub use pixiv::PixivError;
/// The result of fetching a post: canonical URL, HTML caption, raw text,
/// media list and spoiler flag. Produced by [`fetch`].
/// The result of fetching a post: canonical URL, HTML caption, the post's
/// title and body, media list and spoiler flag. Produced by [`fetch`].
#[derive(Debug)]
pub struct Fetched {
/// Canonical URL: x.com/{author}/status/{id} |
/// https://www.pixiv.net/artworks/{id} |
/// https://bsky.app/profile/{handle}/post/{rkey}
/// Canonical URL: `x.com/{author}/status/{id}` |
/// `https://www.pixiv.net/artworks/{id}` |
/// `https://bsky.app/profile/{handle}/post/{rkey}` |
/// `https://www.bilibili.com/opus/{id}`
pub source_url: String,
/// The exact HTML produced by the site's caption().
pub caption: String,
/// Raw post text (tweet text / bsky text / pixiv title).
/// The post's own title, where the platform has one: a pixiv artwork's
/// title, the headline of a bilibili opus post or the title of the video
/// an AV dynamic attaches. Empty on the platforms whose posts are text
/// only (x/twitter, bsky, misskey) and on bilibili posts without a
/// headline.
pub title: String,
/// The post's body text, as the platform exposes it: a tweet, a bsky or
/// misskey post, a bilibili dynamic's text, a pixiv artwork's description
/// (HTML flattened). Empty when the post has no text at all.
pub content: String,
pub media: Vec<crate::media::Media>,
/// Spoiler flag for all media of this post.
pub sensitive: bool,
/// Site id (`"twitter"` / `"bsky"` / `"pixiv"`): the single source of
/// Site id (`"twitter"` / `"bsky"` / `"pixiv"` / `"bilibili"`): the single source of
/// truth for site identity — caption-format lookup, cache-key prefix and
/// the SetFormat whitelist all derive from it. Set by the producing site.
pub site_id: &'static str,
@@ -45,17 +56,38 @@ pub struct Fetched {
pub(crate) _keep_alive: Option<tempfile::TempDir>,
}
/// Pre-escaped values for `{url} {author} {author_url} {title} {tags}`
/// placeholders in user-supplied caption formats.
/// Values for the `{url} {author} {author_url} {title} {content} {tags}`
/// placeholders in user-supplied caption formats, substituted by
/// [`caption_from_fields`] as HTML text (never as an attribute value).
///
/// `author`, `title`, `content` and `tags` come from the site API (post
/// text, display names, descriptions) and are HTML-escaped at construction.
/// `url` and `author_url` stay raw: they are canonical URLs the adapter
/// builds from numeric ids and API-constrained handles/DIDs, so they carry
/// no escapable character — the bot's `/test` report relies on that when it
/// embeds them.
#[derive(Debug)]
pub(crate) struct RenderData {
pub url: String,
pub author: String,
pub author_url: String,
pub title: String,
pub content: String,
pub tags: String,
}
/// The post's text as one string: title and content joined by a line break,
/// each only when it is non-empty. This is what the sites' built-in captions
/// show after the author line, and what the bot quotes when it is long.
pub fn compose_text(title: &str, content: &str) -> String {
match (title.is_empty(), content.is_empty()) {
(false, false) => format!("{title}\n{content}"),
(false, true) => title.to_string(),
(true, false) => content.to_string(),
(true, true) => String::new(),
}
}
impl Fetched {
/// The site this post came from (used for per-site format overrides).
/// A thin alias over [`Fetched::site_id`] kept for callers that read the
@@ -79,21 +111,23 @@ impl Fetched {
&data.author,
&data.author_url,
&data.title,
&data.content,
&data.tags,
),
_ => truncate_caption(&self.caption),
}
}
/// The pre-escaped placeholder values (author, author_url, title, tags)
/// a caller needs to rebuild a caption later, e.g. for a cached post
/// where the [`Fetched`] is no longer available.
pub fn render_fields(&self) -> Option<(&str, &str, &str, &str)> {
/// The pre-escaped placeholder values (author, author_url, title,
/// content, tags) a caller needs to rebuild a caption later, e.g. for a
/// cached post where the [`Fetched`] is no longer available.
pub fn render_fields(&self) -> Option<(&str, &str, &str, &str, &str)> {
self.render_data.as_ref().map(|d| {
(
d.author.as_str(),
d.author_url.as_str(),
d.title.as_str(),
d.content.as_str(),
d.tags.as_str(),
)
})
@@ -139,6 +173,11 @@ pub fn truncate_caption(caption: &str) -> String {
/// [`Fetched::caption_with`]. An empty format returns `built_in` unchanged.
/// The result is truncated to [`MAX_CAPTION_CHARS`] (Telegram's caption
/// limit for HTML parse mode).
///
/// One flat argument per placeholder keeps the two callers (the fresh and the
/// cached caption path) mirroring each other; the same shape as the bot's
/// `debug_report`.
#[allow(clippy::too_many_arguments)]
pub fn caption_from_fields(
format: &str,
built_in: &str,
@@ -146,6 +185,7 @@ pub fn caption_from_fields(
author: &str,
author_url: &str,
title: &str,
content: &str,
tags: &str,
) -> String {
if format.is_empty() {
@@ -158,6 +198,7 @@ pub fn caption_from_fields(
.replace("{author}", author)
.replace("{author_url}", author_url)
.replace("{title}", title)
.replace("{content}", content)
.replace("{tags}", tags),
)
}
@@ -165,7 +206,7 @@ pub fn caption_from_fields(
/// Stable per-post cache key derived from any supported URL, so variant
/// domains (x.com / twitter.com / fxtwitter.com, mobile, `/photo/N`
/// suffixes) map to the same post. Delegates to each registered site's
/// `cache_key` (dispatch order twitter → bsky → pixiv).
/// `cache_key` (in registry order).
pub fn cache_key(url: &str) -> Option<String> {
SITES.iter().find_map(|site| site.cache_key(url))
}
@@ -274,20 +315,21 @@ pub(crate) fn log_once_ffmpeg_missing() {
}
}
/// Site adapter: one impl per supported site (twitter / bsky / pixiv),
/// registered in [`SITES`]. All site-specific knowledge — URL pattern,
/// Site adapter: one impl per supported site (twitter / bsky / misskey /
/// pixiv / bilibili), registered in `SITES`. All site-specific knowledge — URL pattern,
/// cache-key format, fetch, retry policy, media-host headers, startup
/// validation — lives in the site module; the central dispatcher only
/// iterates the registry.
///
/// Async methods return a boxed future (see [`SiteFuture`]): `async fn` /
/// Async methods return a boxed future (see `SiteFuture`): `async fn` /
/// RPITIT in traits are not dyn-compatible (verified on rustc 1.95), and
/// `+ Send` is required since URL/queue workers spawn these futures. The
/// site structs are stateless unit structs, so the boxed futures never
/// borrow from `self` beyond the call's scope.
pub trait Site: Send + Sync {
/// Stable site id (`"twitter"` / `"bsky"` / `"pixiv"`): caption-format
/// lookup, cache-key prefixes and the SetFormat whitelist derive from it.
/// Stable site id (`"twitter"` / `"bsky"` / `"misskey"` / `"pixiv"` /
/// `"bilibili"`): caption-format lookup, cache-key prefixes and the
/// SetFormat whitelist derive from it.
fn id(&self) -> &'static str;
/// URL pattern; the dispatcher's first match wins (dispatch order).
fn pattern(&self) -> &'static Regex;
@@ -323,13 +365,15 @@ pub trait Site: Send + Sync {
type SiteFuture<'a, T, E = FetchError> = Pin<Box<dyn Future<Output = Result<T, E>> + Send + 'a>>;
/// The one registry of supported sites, in dispatch order (twitter → bsky →
/// pixiv). Adding a site = new module + one `Box::new(...)` entry here; the
/// bot crate never lists sites itself.
/// misskey → pixiv → bilibili). Adding a site = new module + one
/// `Box::new(...)` entry here; the bot crate never lists sites itself.
static SITES: LazyLock<Vec<Box<dyn Site>>> = LazyLock::new(|| {
vec![
Box::new(twitter::TwitterSite),
Box::new(bsky::BskySite),
Box::new(misskey::MisskeySite),
Box::new(pixiv::PixivSite),
Box::new(bilibili::BilibiliSite),
]
});
@@ -373,10 +417,24 @@ pub async fn validate_all() -> Vec<(&'static str, String)> {
/// are returned immediately; retrying them only wastes attempts against the
/// source site.
pub async fn fetch(url: &str) -> Result<Option<Fetched>, FetchError> {
fetch_with_attempts(url, MAX_FETCH_ATTEMPTS).await
}
/// [`fetch`] without the retry backoff (one attempt). For callers with a
/// short deadline: an inline query's answer window is measured in seconds, so
/// the 1s + 2s retry sleeps would outlast the query the answer belongs to.
pub async fn fetch_once(url: &str) -> Result<Option<Fetched>, FetchError> {
fetch_with_attempts(url, 1).await
}
/// Total attempts of the retried [`fetch`] (3: the initial try plus two).
const MAX_FETCH_ATTEMPTS: u32 = 3;
async fn fetch_with_attempts(url: &str, attempts: u32) -> Result<Option<Fetched>, FetchError> {
let Some(site) = find_site(url) else {
return Ok(None);
};
for attempt in 0..3u32 {
for attempt in 0..attempts.max(1) {
match site.fetch_from_url(url).await {
Ok(fetched) => {
// Per-request detail: debug only, keyed by the post id.
@@ -389,7 +447,7 @@ pub async fn fetch(url: &str) -> Result<Option<Fetched>, FetchError> {
return Ok(Some(fetched));
}
Err(err) => {
if site.is_retryable(&err) && attempt < 2 {
if site.is_retryable(&err) && attempt + 1 < attempts {
tokio::time::sleep(Duration::from_secs(1 << attempt)).await;
} else {
return Err(err);
@@ -517,6 +575,10 @@ mod tests {
cache_key("https://bsky.app/profile/handle.example/post/3lorem"),
Some("bsky:handle.example/3lorem".into())
);
assert_eq!(
cache_key("https://t.bilibili.com/1245284537985925159"),
Some("bilibili:1245284537985925159".into())
);
assert_eq!(cache_key("https://example.com/not-a-post"), None);
}
@@ -525,16 +587,21 @@ mod tests {
assert_eq!(site_id_from_key("twitter:123"), "twitter");
assert_eq!(site_id_from_key("pixiv:123"), "pixiv");
assert_eq!(site_id_from_key("bsky:handle.example/3lorem"), "bsky");
assert_eq!(site_id_from_key("bilibili:123"), "bilibili");
assert_eq!(site_id_from_key("unknown:1"), "unknown");
assert_eq!(site_id_from_key("no-colon"), "unknown");
}
#[test]
fn registry_lists_all_sites_in_dispatch_order() {
assert_eq!(site_ids(), vec!["twitter", "bsky", "pixiv"]);
assert_eq!(
site_ids(),
vec!["twitter", "bsky", "misskey", "pixiv", "bilibili"]
);
// Enabled sites dispatch; unsupported URLs never match.
assert!(find_site("https://x.com/u/status/1").is_some());
assert!(find_site("https://bsky.app/profile/u/post/3x").is_some());
assert!(find_site("https://misskey.io/notes/abc").is_some());
assert!(find_site("https://t.bilibili.com/1245284537985925159").is_some());
assert!(find_site("https://example.com/x").is_none());
// Cache keys are pattern-driven, independent of the enabled() gate
// (pixiv is disabled in tests without PIXIV_REFRESH_TOKEN).
@@ -562,25 +629,34 @@ mod tests {
// The format string is escaped, the field values are substituted
// verbatim (callers pass the already-escaped render data).
let out = caption_from_fields(
"see {author} at {url} — {title}",
"see {author} at {url} — {title}: {content}",
"",
"https://x.com/u/status/1",
"A &amp; B",
"https://x.com/u",
"hello <world>",
"the body",
"",
);
assert_eq!(
out,
"see A &amp; B at https://x.com/u/status/1 — hello <world>"
"see A &amp; B at https://x.com/u/status/1 — hello <world>: the body"
);
// Empty format keeps the built-in caption untouched.
assert_eq!(
caption_from_fields("", "built-in", "u", "a", "au", "t", "g"),
caption_from_fields("", "built-in", "u", "a", "au", "t", "c", "g"),
"built-in"
);
}
#[test]
fn compose_text_joins_title_and_content() {
assert_eq!(compose_text("标题", "正文"), "标题\n正文");
assert_eq!(compose_text("标题", ""), "标题");
assert_eq!(compose_text("", "正文"), "正文");
assert_eq!(compose_text("", ""), "");
}
#[test]
fn truncate_caption_keeps_short_text() {
assert_eq!(truncate_caption("short"), "short");
@@ -107,10 +107,60 @@ pub fn media_headers(url: &str) -> Option<Vec<(&'static str, String)>> {
}
}
/// Flattens the app API's HTML description into plain text: `<br>` (and `<p>`)
/// become line breaks, other tags are dropped, entities decoded, the ends
/// trimmed. A caption shows text, not markup, so the author's `<a href>` links
/// contribute their link text only.
fn flatten_html(raw: &str) -> String {
let mut out = String::with_capacity(raw.len());
let mut chars = raw.chars().peekable();
while let Some(c) = chars.next() {
// Only `<` followed by `/` or a letter opens a tag — a bare `<` in
// prose ("2 < 3") is text.
let opens_tag = c == '<'
&& chars
.peek()
.is_some_and(|next| *next == '/' || next.is_ascii_alphabetic());
if !opens_tag {
out.push(c);
continue;
}
let mut tag = String::new();
let mut closed = false;
for c in chars.by_ref() {
if c == '>' {
closed = true;
break;
}
tag.push(c);
}
if !closed {
// Unclosed `<…`: keep it as text rather than dropping the tail.
out.push('<');
out.push_str(&tag);
break;
}
// `<br>`, `<br/>`, `<br />` with or without attributes, and both
// halves of a paragraph break the line; everything else is dropped.
let tag = tag
.trim()
.trim_start_matches('/')
.trim_end_matches('/')
.trim()
.to_ascii_lowercase();
if tag == "p" || tag.starts_with("br") {
out.push('\n');
}
}
html_escape::decode_html_entities(&out).trim().to_string()
}
#[derive(Debug)]
pub struct Illustration {
id: String,
title: String,
/// The artwork's description, HTML flattened to plain text.
content: String,
author: String,
author_id: String,
tags: Vec<String>,
@@ -150,6 +200,7 @@ impl Illustration {
pub fn from_model(model: &IllustrationModel) -> Self {
let id = model.id.to_string();
let title = model.title.clone();
let content = flatten_html(&model.caption);
let author = model.user.name.clone();
let author_id = model.user.id.to_string();
let mut tags: Vec<String> = model.tags.iter().map(|tag| tag.name.clone()).collect();
@@ -193,6 +244,7 @@ impl Illustration {
Self {
id,
title,
content,
author,
author_id,
tags,
@@ -218,12 +270,14 @@ impl From<Illustration> for Fetched {
author: encode_text(&illustration.author).into_owned(),
author_url: author_url.clone(),
title: encode_text(&illustration.title).into_owned(),
content: encode_text(&illustration.content).into_owned(),
tags: encode_text(&tags).into_owned(),
});
Fetched {
source_url: url,
caption: illustration.caption(),
title: illustration.title.clone(),
content: illustration.content.clone(),
media: illustration.media,
sensitive: illustration.nsfw,
site_id: "pixiv",
@@ -262,6 +316,7 @@ mod tests {
"illust": {
"id": 123,
"title": "Art <title>",
"caption": "一行说明<br />二行 <a href=\"https://x.example/\">链接</a> &amp; 结尾",
"type": type_,
"image_urls": {
"medium": "medium.jpg",
@@ -284,6 +339,44 @@ mod tests {
Illustration::from_model(&model)
}
/// The description arrives as HTML and becomes plain-text content: breaks
/// kept, tags dropped (links keep their text), entities decoded.
#[test]
fn from_json_maps_description_to_content() {
let v = illust_json("illust", 1, None, Some("o.jpg"), vec![], 0);
let illustration = parse(v);
assert_eq!(illustration.content, "一行说明\n二行 链接 & 结尾");
let fetched: Fetched = illustration.into();
assert_eq!(fetched.title, "Art <title>");
assert_eq!(fetched.content, "一行说明\n二行 链接 & 结尾");
// The built-in caption keeps its layout: the description stays out of
// it and is available through `{content}`.
assert!(!fetched.caption.contains("一行说明"), "{}", fetched.caption);
assert_eq!(
fetched.render_fields().unwrap().3,
"一行说明\n二行 链接 &amp; 结尾"
);
assert!(
fetched
.caption_with("{title}: {content}")
.ends_with("一行说明\n二行 链接 &amp; 结尾")
);
}
#[test]
fn flatten_html_handles_common_markup() {
assert_eq!(flatten_html(""), "");
assert_eq!(flatten_html("plain"), "plain");
assert_eq!(flatten_html("a<br />b<br/>c<br>d"), "a\nb\nc\nd");
// A paragraph break is a blank line, exactly like `<br /><br />` —
// writing it as one newline would flatten the author's paragraphs.
assert_eq!(flatten_html("<p>one</p><p>two</p>"), "one\n\ntwo");
assert_eq!(flatten_html("a &amp; b &lt;c&gt;"), "a & b <c>");
// Nothing to strip: angle brackets that are not a tag survive.
assert_eq!(flatten_html("2 < 3"), "2 < 3");
}
#[test]
fn pattern_matches_all_forms() {
let cases = [
+4
View File
@@ -6,6 +6,10 @@ use serde::Deserialize;
pub struct IllustrationModel {
pub id: u64,
pub title: String,
/// The artwork's description as the app API returns it — HTML in most
/// works (`<br />`, `<a href>`, sometimes `<p>`), empty for many.
#[serde(default)]
pub caption: String,
pub r#type: TypeModel,
pub image_urls: ImageUrlsModel,
pub user: UserInfoModel,
+2 -1
View File
@@ -327,7 +327,8 @@ mod tests {
"https://x.com/nsfw_author/status/2083868672721039569"
);
// The appended media short link (no URL-entity mapping) is stripped.
assert_eq!(fetched.title, "nsfw content");
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "nsfw content");
}
#[test]
+54 -5
View File
@@ -1,7 +1,7 @@
use super::model;
use crate::media::Media;
use crate::site::{FetchError, Fetched, Site, SiteFuture};
use html_escape::{encode_double_quoted_attribute, encode_text};
use html_escape::{decode_html_entities, encode_double_quoted_attribute, encode_text};
use regex::Regex;
use std::sync::LazyLock;
@@ -99,6 +99,7 @@ fn empty_fetched(url: &str) -> Fetched {
// crafted links cannot break the parse (Telegram 400).
caption: encode_text(url).into_owned(),
title: String::new(),
content: String::new(),
media: vec![],
sensitive: true,
site_id: "twitter",
@@ -239,9 +240,17 @@ impl Tweet {
// strip the appended media short link, mirroring FxEmbed's linkFixer
// (no display_text_range arithmetic — see expand_links).
let text = expand_links(&json.text, &json.entities.urls);
// Twitter APIs (syndication AND GraphQL full_text) return the text
// pre-escaped for HTML (`&gt;` `&lt;` `&amp;` `&#39;` …): decode it so
// the stored text is raw. The caption's own escaping then produces
// the rendered form exactly once — without this, `&gt;^ω^&lt;` would
// be double-escaped to `&amp;gt;^ω^&amp;lt;` and the sent message
// would show literal `&gt;^ω^&lt;`.
let text = decode_html_entities(&text).into_owned();
// `name` is the display name, `screen_name` the handle (Python's
// vxtwitter mapping: author = display name, author_id = handle).
let author = json.user.name;
// Display names can carry the same pre-escaped entities.
let author = decode_html_entities(&json.user.name).into_owned();
let author_id = json.user.screen_name;
let mut media = vec![];
for item in json.media_details {
@@ -344,17 +353,20 @@ impl From<Tweet> for Fetched {
fn from(tweet: Tweet) -> Self {
let url = tweet.url();
let author_url = tweet.author_url();
// A tweet has no title: its text is all content.
let render_data = Some(crate::site::RenderData {
url: url.clone(),
author: encode_text(&tweet.author).into_owned(),
author_url: author_url.clone(),
title: encode_text(&tweet.text).into_owned(),
title: String::new(),
content: encode_text(&tweet.text).into_owned(),
tags: String::new(),
});
Fetched {
source_url: url,
caption: tweet.caption(),
title: tweet.text.clone(),
title: String::new(),
content: tweet.text.clone(),
media: tweet.media,
sensitive: tweet.sensitive,
site_id: "twitter",
@@ -409,6 +421,42 @@ mod tests {
}
}
#[test]
fn syndication_text_is_unescaped_before_storing() {
// Real API shape: the text arrives pre-escaped for HTML — e.g. the
// tweet `>^ω^<` comes back as `&gt;^ω^&lt;` (fxtwitter's raw_text for
// 2060196388252827954) and apostrophes as `&#39;`. Storing it raw and
// escaping once at caption build avoids the double-escape that would
// show literal `&gt;`/`&lt;`/`&amp;` in the sent message.
let raw = serde_json::json!({
"__typename": "Tweet",
"id_str": "1",
"text": "&gt;^ω^&lt; &amp; more &#39;quoted&#39; https://t.co/abc123",
"user": { "name": "O&#39;Brien", "screen_name": "h" },
"entities": { "urls": [] },
"mediaDetails": []
});
let tweet = Tweet::from_syndication_json(&raw.to_string()).unwrap();
// The appended media short link is stripped, then entities decoded.
assert_eq!(tweet.text, ">^ω^< & more 'quoted'");
assert_eq!(tweet.author, "O'Brien");
let fetched: Fetched = tweet.into();
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, ">^ω^< & more 'quoted'");
// The caption escapes the raw text exactly once (encode_text covers
// & < >; apostrophes stay literal — they are harmless in text).
assert!(
fetched.caption.contains("&gt;^ω^&lt; &amp; more 'quoted'"),
"caption: {}",
fetched.caption
);
assert!(
!fetched.caption.contains("&amp;gt;"),
"double-escaped text: {}",
fetched.caption
);
}
#[test]
fn cache_key_prefixes_tweet_id() {
assert_eq!(
@@ -453,7 +501,8 @@ mod tests {
fetched.source_url,
"https://x.com/author_handle/status/861627479294746624"
);
assert_eq!(fetched.title, "a & b <c>");
assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "a & b <c>");
assert!(fetched.sensitive);
assert_eq!(fetched.media.len(), 2);
match &fetched.media[0] {
+7 -4
View File
@@ -1,11 +1,11 @@
[package]
name = "xmedia-bot"
version = "1.2.2"
version = "1.7.0"
edition = "2024"
[dependencies]
teloxide = { version = "0.17", default-features = false, features = ["webhooks-axum", "macros", "rustls", "ctrlc_handler"] }
tokio = { version = "1.40", features = ["rt-multi-thread", "macros"] }
tokio = { version = "1.40", features = ["rt-multi-thread", "macros", "time", "sync"] }
serde = { version = "1", features = ["derive"] }
serde_json = "1"
log = "0.4"
@@ -13,8 +13,8 @@ pretty_env_logger = "0.5"
dotenv = "0.15"
url = "2.5.2"
html-escape = "0.2"
rusqlite = { version = "0.32", features = ["bundled"] }
rand = "0.8"
rusqlite = { version = "0.40", features = ["bundled"] }
rand = "0.10"
tempfile = "3"
parking_lot = "0.12"
bytes = "1"
@@ -23,3 +23,6 @@ zune-jpeg = "0.5"
fast_image_resize = "6"
jpeg-encoder = "0.7"
x-media = { path = "../x-media" }
[dev-dependencies]
tokio = { version = "1.40", features = ["test-util"] }
+8 -1
View File
@@ -1,5 +1,6 @@
//! Central env handling. The only other places that read env are
//! `Bot::from_env` (TELOXIDE_TOKEN) and x-media (PIXIV_REFRESH_TOKEN).
//! `Bot::from_env` (TELOXIDE_TOKEN) and x-media (PIXIV_REFRESH_TOKEN,
//! TWITTER_AUTH_TOKEN, BILIBILI_COOKIE).
use std::env;
use std::net::IpAddr;
@@ -12,6 +13,10 @@ pub struct Config {
pub edit_message_ttl: Duration,
/// LINK_CACHE_TTL_SECONDS, default 604800 (7 days).
pub link_cache_ttl: Duration,
/// CAPTION_QUOTE_TEXT_CHARS, default 200: a post whose text (title plus
/// content) is at least this many characters gets that text wrapped in an
/// expandable blockquote inside its caption. `0` disables the wrap.
pub caption_quote_text_chars: usize,
// Webhook settings (moved out of main; names/defaults unchanged).
pub webhook_enabled: bool,
pub webhook_url: Option<url::Url>,
@@ -57,6 +62,7 @@ impl Config {
Duration::from_secs(parse_u64("EDIT_MESSAGE_TTL_SECONDS", 24 * 3600));
let link_cache_ttl =
Duration::from_secs(parse_u64("LINK_CACHE_TTL_SECONDS", 7 * 24 * 3600));
let caption_quote_text_chars = parse_u64("CAPTION_QUOTE_TEXT_CHARS", 200) as usize;
let webhook_enabled = env::var("WEBHOOK")
.is_ok_and(|v| matches!(v.to_lowercase().as_str(), "true" | "yes" | "1"));
@@ -92,6 +98,7 @@ impl Config {
admin_ids,
edit_message_ttl,
link_cache_ttl,
caption_quote_text_chars,
webhook_enabled,
webhook_url,
webhook_listen,
+124
View File
@@ -0,0 +1,124 @@
//! Runtime context: the collaborators a handler needs, injected as one struct
//! so tests can substitute a scripted sender and tempdir-backed stores.
//!
//! The production context is assembled from the process-wide statics
//! ([`AppContext::from_statics`]); the spawned worker closures hold
//! [`CONTEXT`], which is `'static` for that reason.
use crate::config::Config;
use crate::handlers::{CHAT_STORE, CONFIG, LINK_CACHE, TASK_QUEUE};
use crate::link_cache::LinkCache;
use crate::media_sender::MediaSender;
use crate::queue::PersistentTaskQueue;
use crate::send::BOT;
use crate::state::ChatStore;
use std::sync::LazyLock;
pub struct AppContext<'a> {
pub sender: &'a dyn MediaSender,
pub chat_store: &'a ChatStore,
pub task_queue: &'a PersistentTaskQueue,
pub link_cache: &'a LinkCache,
pub config: &'a Config,
}
impl<'a> AppContext<'a> {
/// The stores are the process-wide statics; `sender` is whatever the caller
/// was handed (the dispatcher's `Bot` clone for update handlers, the shared
/// queue `Bot` for the worker loops). Update handlers build their own
/// context from the `Bot` they received so the same code path works with an
/// injected mock in tests.
pub fn from_statics(sender: &'a dyn MediaSender) -> AppContext<'a> {
AppContext {
sender,
chat_store: &CHAT_STORE,
task_queue: &TASK_QUEUE,
link_cache: &LINK_CACHE,
config: &CONFIG,
}
}
}
/// The URL/queue workers' context: `'static` because `tokio::spawn`ed closures
/// and the queue's handler type require it.
pub static CONTEXT: LazyLock<AppContext<'static>> =
LazyLock::new(|| AppContext::from_statics(&*BOT));
/// Test support: a tempdir-backed set of stores plus the context borrowing
/// them, so a handler test needs one line of setup.
#[cfg(test)]
pub(crate) mod test_support {
use super::*;
use std::sync::Arc;
pub(crate) struct TestStores {
_dir: tempfile::TempDir,
pool: Arc<crate::db::DbPool>,
chat_store: ChatStore,
task_queue: PersistentTaskQueue,
link_cache: LinkCache,
config: Config,
}
impl TestStores {
pub(crate) fn new() -> Self {
let dir = tempfile::tempdir().unwrap();
let pool = crate::db::open_store(dir.path().join("ctx.db").to_str().unwrap()).unwrap();
TestStores {
_dir: dir,
chat_store: ChatStore::new(Arc::clone(&pool)),
task_queue: PersistentTaskQueue::new(Arc::clone(&pool)),
link_cache: LinkCache::new(Arc::clone(&pool)),
config: Config::load(),
pool,
}
}
pub(crate) fn ctx<'a>(&'a self, sender: &'a dyn MediaSender) -> AppContext<'a> {
AppContext {
sender,
chat_store: &self.chat_store,
task_queue: &self.task_queue,
link_cache: &self.link_cache,
config: &self.config,
}
}
pub(crate) fn chat_store(&self) -> &ChatStore {
&self.chat_store
}
/// The parsed config, mutable so a test can pin a knob (e.g. the
/// caption-quote threshold) instead of depending on the environment.
pub(crate) fn config_mut(&mut self) -> &mut Config {
&mut self.config
}
pub(crate) fn link_cache(&self) -> &LinkCache {
&self.link_cache
}
/// Rows persisted in the task queue: what "queued for retry" looks like
/// from the outside.
pub(crate) async fn queued_tasks(&self) -> i64 {
let pool = Arc::clone(&self.pool);
pool.with_conn(|conn| {
conn.query_row("SELECT COUNT(*) FROM tasks", [], |row| row.get(0))
})
.await
.unwrap()
}
/// The single queued task payload, for asserting what was rescheduled.
pub(crate) async fn queued_payload(&self) -> serde_json::Value {
let pool = Arc::clone(&self.pool);
let payload: String = pool
.with_conn(|conn| {
conn.query_row("SELECT payload FROM tasks LIMIT 1", [], |row| row.get(0))
})
.await
.unwrap();
serde_json::from_str(&payload).unwrap()
}
}
}
+12
View File
@@ -134,6 +134,12 @@ fn rusqlite_error(e: std::io::Error) -> rusqlite::Error {
/// Creates the `tasks`, `chat_state` and `link_cache` tables (idempotent).
/// The three stores used to own their own schema; keeping it in one place
/// means one initialization for the whole database file.
///
/// ⚠️ Schema-change reminder (deferred, see `docs/architecture-refactor.md`
/// §5): this is a plain `CREATE TABLE IF NOT EXISTS` with no versioning.
/// Before any column/table change that must migrate existing databases, land
/// the `PRAGMA user_version` migration chain first (`MIGRATIONS: &[&str]` +
/// `migrate(conn)`), then restructure this function.
pub fn schema_init(conn: &Connection) -> rusqlite::Result<()> {
conn.execute_batch(
"CREATE TABLE IF NOT EXISTS tasks (id TEXT PRIMARY KEY, payload TEXT NOT NULL, \
@@ -154,3 +160,9 @@ pub fn now_f64() -> f64 {
.map(|d| d.as_secs_f64())
.unwrap_or(0.0)
}
/// Unix timestamp in whole seconds. Same clock as [`now_f64`], for fields
/// that store integer seconds (chat-state expiry, edit prompts).
pub fn unix_now() -> i64 {
now_f64() as i64
}
File diff suppressed because it is too large Load Diff
+321
View File
@@ -0,0 +1,321 @@
//! Callback query handling: the edit-before-forward prompt's `"forward"` and
//! `"template|<name>"` buttons.
//!
//! [`callback_query_handler`] is the dptree entry; it only pulls the plain
//! values out of the teloxide update and hands them to [`handle_callback`],
//! which holds the button logic and is driven directly by tests.
use crate::ctx::AppContext;
use crate::db::unix_now;
use crate::send::{self, Task};
use teloxide::RequestError;
use teloxide::prelude::*;
use teloxide::types::{CallbackQuery, CallbackQueryId, MessageId};
/// The `"forward"` button's data.
const FORWARD: &str = "forward";
/// Prefix of a template button's data: `"template|<name>"`.
const TEMPLATE_PREFIX: &str = "template|";
pub async fn callback_query_handler(bot: Bot, query: CallbackQuery) -> Result<(), RequestError> {
let Some(message) = &query.message else {
return respond(());
};
let Some(data) = query.data.clone() else {
return respond(());
};
let ctx = AppContext::from_statics(&bot);
handle_callback(
&ctx,
query.id.clone(),
message.chat().id.0,
message.id().0 as i64,
&data,
)
.await;
respond(())
}
/// Handles one button press on the edit-before-forward prompt.
async fn handle_callback(
ctx: &AppContext<'_>,
callback_query_id: CallbackQueryId,
chat_id: i64,
prompt_message_id: i64,
data: &str,
) {
let ttl_secs = ctx.config.edit_message_ttl.as_secs() as i64;
let chat_data = ctx.chat_store.get(chat_id).await;
let edit = chat_data.edit_message.get(&prompt_message_id).cloned();
let Some(edit) = edit else {
log::debug!("callback from {chat_id}: no edit record for prompt {prompt_message_id}");
let _ = ctx
.sender
.answer_callback_query(callback_query_id, Some("Expired".to_string()))
.await;
return;
};
// Lazy expiry: a stale record (past the TTL, not yet swept) is dropped.
if edit.created_at + ttl_secs <= unix_now() {
ctx.chat_store
.update(chat_id, |data| {
data.edit_message.remove(&prompt_message_id);
})
.await;
let _ = ctx
.sender
.answer_callback_query(callback_query_id, Some("Expired".to_string()))
.await;
return;
}
log::info!("callback from {chat_id} on prompt {prompt_message_id}: {data}");
if data == FORWARD {
match chat_data.forward_channel_id {
Some(channel_id) => {
let forward_task = Task::ForwardMessages {
from_chat_id: edit.chat_id,
to_chat_id: channel_id,
message_ids: edit.forward_message_ids.clone(),
notify_chat_id: Some(chat_id),
notify_message_id: Some(prompt_message_id),
};
let (answer, settled) = match send::forward_messages(ctx, &forward_task).await {
Ok(()) => {
log::info!(
"forwarded {} message(s) to channel {channel_id}",
edit.forward_message_ids.len()
);
("✅ Forwarded".to_string(), true)
}
Err(send::SendError::Retryable {
delay_seconds,
task,
}) => {
log::info!("forward queued for retry in {delay_seconds:.1}s");
send::enqueue_retry(ctx.task_queue, *task, delay_seconds).await;
("Forward queued for retry.".to_string(), false)
}
Err(send::SendError::Permanent { message, .. }) => {
log::error!("forward failed permanently: {message}");
(format!("Forward failed: {message}"), false)
}
};
if settled {
// The prompt is done: drop it and its record.
let _ = ctx
.sender
.delete_message(ChatId(chat_id), MessageId(prompt_message_id as i32))
.await;
ctx.chat_store
.update(chat_id, |data| {
data.edit_message.remove(&prompt_message_id);
})
.await;
}
let _ = ctx
.sender
.answer_callback_query(callback_query_id, Some(answer))
.await;
}
None => {
log::debug!("forward callback without a forward channel set");
let _ = ctx
.sender
.answer_callback_query(
callback_query_id,
Some("No forward channel set.".to_string()),
)
.await;
}
}
return;
}
if let Some(name) = data.strip_prefix(TEMPLATE_PREFIX) {
if let Some(template_html) = chat_data.template.get(name).cloned()
&& let Some(first_forward_id) = edit.forward_message_ids.first().copied()
{
// Raw template including the [] placeholder (Python parity).
let _ = ctx
.sender
.edit_message_caption(
ChatId(chat_id),
MessageId(first_forward_id as i32),
template_html,
)
.await;
ctx.chat_store
.update(chat_id, |data| {
if let Some(entry) = data.edit_message.get_mut(&prompt_message_id) {
entry.template = name.to_string();
}
})
.await;
log::info!("template '{name}' applied to prompt {prompt_message_id}");
}
let _ = ctx
.sender
.answer_callback_query(callback_query_id, None)
.await;
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::ctx::test_support::TestStores;
use crate::media_sender::test_support::{MockSender, Outcome};
use crate::state::EditMessage;
use teloxide::ApiError;
/// The edit-before-forward prompt's message id in these tests.
const PROMPT_ID: i64 = 7;
/// The message the prompt refers to (the one whose caption is swapped).
const FORWARDED_ID: i64 = 9;
fn api_error() -> RequestError {
RequestError::Api(ApiError::Unknown("Bad Request: chat not found".into()))
}
fn callback_id() -> CallbackQueryId {
CallbackQueryId("cb-1".to_string())
}
/// Seeds a live prompt record plus a forward channel and a template;
/// `created_at` backdates the record for the expiry cases.
async fn seed_prompt(ctx: &AppContext<'_>, created_at: i64) {
ctx.chat_store
.update(1, |data| {
data.forward_channel_id = Some(2);
data.template
.insert("tpl".to_string(), "<b>[]</b>".to_string());
data.edit_message.insert(
PROMPT_ID,
EditMessage {
url: "https://x.com/u/status/1".into(),
chat_id: 1,
forward_message_ids: vec![FORWARDED_ID],
template: String::new(),
created_at,
},
);
})
.await;
}
#[tokio::test]
async fn template_button_swaps_the_caption_and_records_the_choice() {
let sender = MockSender::scripted(vec![Outcome::EditOk], api_error);
let stores = TestStores::new();
let ctx = stores.ctx(&sender);
seed_prompt(&ctx, crate::db::unix_now()).await;
handle_callback(&ctx, callback_id(), 1, PROMPT_ID, "template|tpl").await;
assert_eq!(
sender.calls(),
vec!["edit_message_caption", "answer_callback_query"]
);
// The raw template, including the [] the user edits into.
assert_eq!(sender.captions(), vec!["<b>[]</b>"]);
assert_eq!(sender.answers(), vec![None]);
let data = ctx.chat_store.get(1).await;
assert_eq!(data.edit_message[&PROMPT_ID].template, "tpl");
}
#[tokio::test]
async fn forward_button_copies_then_clears_the_prompt() {
let sender = MockSender::scripted(vec![Outcome::CopyOk], api_error);
let stores = TestStores::new();
let ctx = stores.ctx(&sender);
seed_prompt(&ctx, crate::db::unix_now()).await;
handle_callback(&ctx, callback_id(), 1, PROMPT_ID, "forward").await;
assert_eq!(
sender.calls(),
vec!["copy_messages", "delete_message", "answer_callback_query"]
);
assert_eq!(sender.answers(), vec![Some("✅ Forwarded".to_string())]);
assert!(
ctx.chat_store.get(1).await.edit_message.is_empty(),
"a settled prompt must drop its record"
);
}
#[tokio::test]
async fn forward_without_a_channel_is_reported() {
let sender = MockSender::scripted(vec![], api_error);
let stores = TestStores::new();
let ctx = stores.ctx(&sender);
seed_prompt(&ctx, crate::db::unix_now()).await;
ctx.chat_store
.update(1, |data| data.forward_channel_id = None)
.await;
handle_callback(&ctx, callback_id(), 1, PROMPT_ID, "forward").await;
assert_eq!(sender.calls(), vec!["answer_callback_query"]);
assert_eq!(
sender.answers(),
vec![Some("No forward channel set.".to_string())]
);
}
#[tokio::test]
async fn retryable_forward_is_queued_and_keeps_the_prompt() {
use teloxide::types::Seconds;
let sender = MockSender::scripted(vec![Outcome::CopyErr], || {
RequestError::RetryAfter(Seconds::from_seconds(7))
});
let stores = TestStores::new();
let ctx = stores.ctx(&sender);
seed_prompt(&ctx, crate::db::unix_now()).await;
handle_callback(&ctx, callback_id(), 1, PROMPT_ID, "forward").await;
assert_eq!(
sender.calls(),
vec!["copy_messages", "answer_callback_query"]
);
assert_eq!(
sender.answers(),
vec![Some("Forward queued for retry.".to_string())]
);
assert_eq!(stores.queued_tasks().await, 1);
// The prompt is not settled: the queued retry still needs the record.
assert!(
ctx.chat_store
.get(1)
.await
.edit_message
.contains_key(&PROMPT_ID)
);
}
#[tokio::test]
async fn unknown_and_expired_prompts_answer_expired() {
let sender = MockSender::scripted(vec![], api_error);
let stores = TestStores::new();
let ctx = stores.ctx(&sender);
// No record at all.
handle_callback(&ctx, callback_id(), 1, PROMPT_ID, "forward").await;
assert_eq!(sender.answers(), vec![Some("Expired".to_string())]);
// A record past its TTL (nothing swept it yet) is dropped on use.
let stale = crate::db::unix_now() - ctx.config.edit_message_ttl.as_secs() as i64 - 1;
seed_prompt(&ctx, stale).await;
handle_callback(&ctx, callback_id(), 1, PROMPT_ID, "forward").await;
assert_eq!(
sender.answers(),
vec![Some("Expired".to_string()), Some("Expired".to_string())]
);
assert!(
ctx.chat_store.get(1).await.edit_message.is_empty(),
"the expired record must be dropped"
);
assert_eq!(sender.calls(), vec!["answer_callback_query"; 2]);
}
}
+660
View File
@@ -0,0 +1,660 @@
//! Bot command parsing, the `/`-command executor and `setMyCommands`
//! registration. URL/inline/callback flows live in their own modules.
use super::urls::{PostSend, url_media};
use super::{CHAT_STORE, CONFIG, LINK_CACHE, log_key, reply, reply_html};
use crate::ctx::AppContext;
use teloxide::RequestError;
use teloxide::prelude::*;
use teloxide::types::{ChatId, Message, Recipient};
use teloxide::utils::command::{BotCommands, ParseError};
#[derive(BotCommands, Clone)]
#[command(
rename_rule = "snake_case",
description = "Turn X/Pixiv/Bluesky links into media messages"
)]
pub(crate) enum Command {
#[command(description = "Get started")]
Start,
#[command(description = "Show command help")]
Help,
#[command(
description = "Set forward channel (@channel or ID)",
parse_with = "split"
)]
SetForwardChannel(String),
#[command(description = "Remove forward channel")]
RemoveForwardChannel,
#[command(description = "Toggle edit-before-forward")]
EditBeforeForward,
#[command(
description = "Reply with [] to save as template",
parse_with = "split"
)]
SetTemplate(String),
#[command(description = "Show chat state (debug; admin only)")]
BotDict,
#[command(description = "Set site caption format", parse_with = "split")]
SetFormat(String),
#[command(
description = "Clear link cache (admin; optional URL, else all)",
parse_with = "split"
)]
ClearCache(String),
#[command(
description = "Send a link's media (no forwarding)",
parse_with = parse_arg_remainder
)]
Test(String),
#[command(
description = "Parse a link and report it (debug; nothing sent)",
parse_with = parse_arg_remainder
)]
Debug(String),
}
/// `/test` and `/debug` argument parser: the whole remainder after the command
/// name, trimmed. The built-in `split` parser takes exactly one space-separated
/// token and rejects the rest, so a URL followed by a trailing space (or
/// pasted text) would silently fall through to the URL flow instead.
fn parse_arg_remainder(s: String) -> Result<(String,), ParseError> {
Ok((s.trim().to_string(),))
}
enum SetForwardChannelError {
EmptyParameter,
NotChannel,
NotAdmin,
NotBotAdmin(RequestError),
NotBotCanPost,
}
async fn set_forward_channel_handler(
bot: &Bot,
message: &Message,
channel: String,
) -> Result<i64, SetForwardChannelError> {
if channel.is_empty() {
return Err(SetForwardChannelError::EmptyParameter);
}
let channel = match channel.parse::<i64>() {
Ok(id) => Recipient::Id(ChatId(id)),
Err(_) => Recipient::ChannelUsername(channel),
};
if let Some(from) = &message.from {
log::info!(
"Set forward channel for {} ({}) to {}",
from.full_name(),
message.chat.id,
channel
);
}
let chat = match bot.get_chat(channel.clone()).await {
Err(e) => {
log::error!("Failed to get channel {}: {}", channel, e);
return Err(SetForwardChannelError::NotBotAdmin(e));
}
Ok(chat) => chat,
};
if !chat.is_channel() {
return Err(SetForwardChannelError::NotChannel);
}
let channel_id = chat.id.0;
// The sender must be a channel administrator. Compare against the
// sender's user id, NOT the chat id (they only coincide in private
// chats, so the old check broke group usage).
let Some(sender) = message.from.as_ref() else {
return Err(SetForwardChannelError::NotAdmin);
};
match bot.get_chat_administrators(channel.clone()).await {
Err(e) => {
log::error!("Failed to get channel administrators {}: {}", channel, e);
return Err(SetForwardChannelError::NotBotAdmin(e));
}
Ok(admins) => {
if !admins.iter().any(|admin| admin.user.id == sender.id) {
return Err(SetForwardChannelError::NotAdmin);
}
// The bot itself must be an admin that can post; a missing
// bot entry must not pass silently (copy would fail later).
let bot_id = match bot.get_me().await {
Ok(me) => me.user.id,
Err(e) => return Err(SetForwardChannelError::NotBotAdmin(e)),
};
let bot_ok = admins
.iter()
.any(|admin| admin.user.id == bot_id && admin.can_post_messages());
if !bot_ok {
return Err(SetForwardChannelError::NotBotCanPost);
}
}
}
Ok(channel_id)
}
pub(crate) async fn execute_command(
bot: &Bot,
message: &Message,
command: Command,
) -> Result<(), RequestError> {
match command {
Command::Start => {
bot.send_message(message.chat.id, "Hello!").await?;
}
Command::Help => {
bot.send_message(message.chat.id, Command::descriptions().to_string())
.await?;
}
Command::SetForwardChannel(channel) => {
let result = match set_forward_channel_handler(bot, message, channel).await {
Ok(channel_id) => {
CHAT_STORE
.update(message.chat.id.0, |data| {
data.forward_channel_id = Some(channel_id);
})
.await;
"Add successfully.".to_string()
}
Err(SetForwardChannelError::EmptyParameter) => {
"Receive empty parameter.\nYou should enter a channel id or username"
.to_string()
}
Err(SetForwardChannelError::NotChannel) => {
"Given id / username is not a channel".to_string()
}
Err(SetForwardChannelError::NotAdmin) => {
"You are not an administrator of the channel".to_string()
}
Err(SetForwardChannelError::NotBotAdmin(e)) => {
e.to_string() + "\nPlease add the bot as an admin to the channel"
}
Err(SetForwardChannelError::NotBotCanPost) => {
"Bot can't post messages to the channel".to_string()
}
};
reply(bot, message.chat.id.0, message.id, result).await?;
}
Command::RemoveForwardChannel => {
let chat_id = message.chat.id.0;
let text = CHAT_STORE
.update(chat_id, |data| {
if data.forward_channel_id.is_some() {
data.forward_channel_id = None;
"Remove successfully.".to_string()
} else {
"No channel to remove.".to_string()
}
})
.await;
reply(bot, message.chat.id.0, message.id, text).await?;
}
Command::EditBeforeForward => {
let chat_id = message.chat.id.0;
let text = CHAT_STORE
.update(chat_id, |data| {
if data.forward_channel_id.is_none() {
"Please enable forward channel first.".to_string()
} else if data.edit_before_forward {
data.edit_before_forward = false;
data.edit_message.clear();
"Disable edit before forward.".to_string()
} else {
data.edit_before_forward = true;
"Enable edit before forward.".to_string()
}
})
.await;
reply(bot, message.chat.id.0, message.id, text).await?;
}
Command::SetTemplate(name) => {
let chat_id = message.chat.id.0;
let text = match message.reply_to_message() {
None => "Please reply to a message to set as template.".to_string(),
Some(reply) => {
let reply_text = reply.text().unwrap_or_default();
if !reply_text.contains("[]") {
"Please reply to a message with [] to set as template.".to_string()
} else if name.is_empty() {
"Please provide a name for the template.".to_string()
} else {
CHAT_STORE
.update(chat_id, |data| {
data.template.insert(
name,
html_escape::encode_text(reply_text).into_owned(),
);
})
.await;
"Template set.".to_string()
}
}
};
reply(bot, message.chat.id.0, message.id, text).await?;
}
Command::BotDict => {
// Debug dump of the chat's persisted state: admin only (it echoes
// forward-channel ids and templates to whoever asks).
let sender_id = message
.from
.as_ref()
.map(|user| user.id.0 as i64)
.unwrap_or(-1);
if !CONFIG.admin_ids.contains(&sender_id) {
reply(bot, message.chat.id.0, message.id, "Admin only.").await?;
return Ok(());
}
let chat_data = CHAT_STORE.get(message.chat.id.0).await;
let debug = html_escape::encode_text(&format!("{chat_data:?}")).into_owned();
// A chat with many templates/edit records exceeds Telegram's 4096
// char message limit; the dump is plain text (no parse mode), so a
// plain byte-boundary cut is safe.
let end = debug.floor_char_boundary(MAX_DEBUG_DUMP_CHARS.min(debug.len()));
let text = if end < debug.len() {
format!("{}", &debug[..end])
} else {
debug
};
reply(bot, message.chat.id.0, message.id, text).await?;
}
Command::SetFormat(arg) => {
let chat_id = message.chat.id.0;
let (site, format) = match arg.split_once(char::is_whitespace) {
Some((site, format)) if !format.trim().is_empty() => {
(site.trim(), format.trim().to_string())
}
_ => {
reply(
bot,
message.chat.id.0,
message.id,
"Usage: /set_format <site> <format>",
)
.await?;
return Ok(());
}
};
if !x_media::site::site_ids().contains(&site) {
reply(
bot,
message.chat.id.0,
message.id,
"Unknown site. Use twitter, bsky, pixiv, misskey or bilibili.",
)
.await?;
return Ok(());
}
CHAT_STORE
.update(chat_id, |data| {
data.message_format.insert(site.to_string(), format);
})
.await;
reply(bot, message.chat.id.0, message.id, "Format set.").await?;
}
Command::ClearCache(arg) => {
let sender_id = message
.from
.as_ref()
.map(|user| user.id.0 as i64)
.unwrap_or(-1);
if !CONFIG.admin_ids.contains(&sender_id) {
reply(bot, message.chat.id.0, message.id, "Admin only.").await?;
return Ok(());
}
let arg = arg.trim();
if arg.is_empty() {
let removed = LINK_CACHE.clear(None).await;
log::info!("cache cleared by {sender_id}: {removed} entries");
reply(
bot,
message.chat.id.0,
message.id,
format!("Cleared {removed} cached entr{}.", plural(removed)),
)
.await?;
} else {
let key = match x_media::site::cache_key(arg) {
Some(key) => key,
None => {
reply(
bot,
message.chat.id.0,
message.id,
"Unrecognized link. Use a twitter/x, pixiv, bsky, misskey or bilibili post URL.",
)
.await?;
return Ok(());
}
};
let removed = LINK_CACHE.clear(Some(&key)).await;
log::info!("cache entry cleared by {sender_id}: {key} ({removed} rows)");
reply(
bot,
message.chat.id.0,
message.id,
format!(
"Cleared cache for {arg} ({} entr{}).",
removed,
plural(removed)
),
)
.await?;
}
}
Command::Test(arg) => {
let url = arg.trim();
if url.is_empty() {
reply(
bot,
message.chat.id.0,
message.id,
"Usage: /test <post url>",
)
.await?;
return Ok(());
}
if x_media::site::cache_key(url).is_none() {
reply(
bot,
message.chat.id.0,
message.id,
"No enabled site matches this link (twitter/x, pixiv, bsky, misskey or bilibili).",
)
.await?;
return Ok(());
}
// The ordinary link pipeline with the chat's post-send actions
// suppressed: the media is sent (and cached) like a normal link,
// but nothing is forwarded to the channel and no
// edit-before-forward prompt opens. Info level echoes the
// normalized key (never the raw URL) per the logging convention.
log::info!("test: sending [key={}]", log_key(url));
let ctx = AppContext::from_statics(bot);
url_media(
&ctx,
message.chat.id.0,
message.id.0 as i64,
url,
PostSend::Suppressed,
)
.await;
}
Command::Debug(arg) => {
let url = arg.trim();
if url.is_empty() {
reply(
bot,
message.chat.id.0,
message.id,
"Usage: /debug <post url>",
)
.await?;
return Ok(());
}
// Debug tool: report the parse result only — nothing is sent,
// cached or forwarded.
log::info!("debug: parsing [key={}]", log_key(url));
match x_media::site::fetch(url).await {
Ok(None) => {
reply(
bot,
message.chat.id.0,
message.id,
"No enabled site matches this link (twitter/x, pixiv, bsky, misskey or bilibili).",
)
.await?;
}
Err(e) => {
reply(
bot,
message.chat.id.0,
message.id,
format!("Fetch failed: {e}"),
)
.await?;
}
Ok(Some(fetched)) => {
let report = debug_report(
url,
fetched.site_name(),
&fetched.source_url,
&fetched.title,
&fetched.content,
fetched.render_fields(),
fetched.sensitive,
&fetched.caption,
&fetched.media,
);
// HTML report: the caption renders inside a <blockquote>
// exactly as it will appear in the sent media message.
reply_html(bot, message.chat.id.0, message.id, report).await?;
}
}
}
}
Ok(())
}
/// `"y"` for one, `"ies"` for anything else — "1 entry" / "2 entries".
fn plural(n: usize) -> &'static str {
if n == 1 { "y" } else { "ies" }
}
/// Registers the bot's command list with Telegram so clients show it in the
/// `/` menu (Bot API `setMyCommands`).
pub async fn register_commands(bot: &Bot) -> Result<(), RequestError> {
let commands = Command::bot_commands();
bot.set_my_commands(commands.clone()).await?;
log::info!("registered {} commands", commands.len());
Ok(())
}
/// Telegram's plain-text message limit is 4096 chars; the report stays under
/// it even for very large threads (many media lines + a long caption).
const MAX_DEBUG_REPORT_CHARS: usize = 4000;
/// Cap for the `/bot_dict` debug dump: the state is echoed as one plain-text
/// message, so it must stay under Telegram's 4096-char limit.
const MAX_DEBUG_DUMP_CHARS: usize = 3500;
/// Builds the HTML report for the `/debug` command: what the parser produced
/// for a link (site, canonical URL, title/author/tags, caption and the media
/// list) — no media is sent and nothing is cached or forwarded. Sent with
/// HTML parse mode: raw fields are escaped, the pre-escaped render fields are
/// embedded as-is, and the caption is wrapped in a `<blockquote>` so it shows
/// exactly as it will render in the sent media message. Fields are passed
/// individually so the formatter stays a pure function testable without
/// constructing a `Fetched` (its render fields are `pub(crate)` to the
/// x-media crate).
#[allow(clippy::too_many_arguments)]
fn debug_report(
url: &str,
site_id: &str,
source_url: &str,
title: &str,
content: &str,
render: Option<(&str, &str, &str, &str, &str)>,
sensitive: bool,
caption: &str,
media: &[x_media::media::Media],
) -> String {
let mut lines = vec![
format!("Parse result for {}", html_escape::encode_text(url)),
format!("site: {site_id}"),
format!(
"key: {}",
html_escape::encode_text(
&x_media::site::cache_key(url).unwrap_or_else(|| "<unsupported>".to_string())
)
),
];
lines.push(format!(
"source_url: {}",
html_escape::encode_text(source_url)
));
lines.push(format!("title: {}", html_escape::encode_text(title)));
lines.push(format!("content: {}", html_escape::encode_text(content)));
if let Some((author, author_url, _title, _content, tags)) = render {
// The render fields are already pre-escaped for HTML captions; embed
// them as-is so the report renders them exactly like the final
// caption. `author_url` is raw and gets escaped here.
lines.push(format!("author: {author}"));
lines.push(format!(
"author_url: {}",
html_escape::encode_text(author_url)
));
lines.push(format!("tags: {tags}"));
}
lines.push(format!("sensitive: {sensitive}"));
// The caption is wrapped in a <blockquote> so the report (an HTML
// message) shows it exactly as it will render in the sent media caption
// — escaped text and links included.
lines.push(format!(
"caption: <blockquote>{}</blockquote>",
x_media::site::truncate_caption(caption)
));
lines.push(format!("media ({}):", media.len()));
for (i, item) in media.iter().enumerate() {
let kind = match item {
x_media::media::Media::Illustration { .. } => "image",
x_media::media::Media::Video { .. } => "video",
x_media::media::Media::Animated { .. } => "gif",
};
lines.push(format!(
" {}. {kind}: {}",
i + 1,
html_escape::encode_text(item.url())
));
}
let mut out = lines.join(
"
",
);
if out.chars().count() > MAX_DEBUG_REPORT_CHARS {
let end = out.floor_char_boundary(MAX_DEBUG_REPORT_CHARS - 1);
out = format!("{}", &out[..end]);
}
out
}
#[cfg(test)]
mod tests {
use super::{MAX_DEBUG_REPORT_CHARS, debug_report};
use x_media::media::Media;
#[test]
fn debug_report_renders_fields_and_media() {
let media = vec![
Media::Illustration {
title: None,
url: "https://cdn.example/1.jpg".into(),
thumbnail_url: None,
fallback_url: None,
},
Media::Video {
title: None,
url: "https://cdn.example/2.mp4".into(),
thumbnail_url: "https://cdn.example/2.jpg".into(),
},
];
let report = debug_report(
"https://x.com/u/status/1",
"twitter",
"https://x.com/u/status/1",
"My title",
"My content",
Some((
"Author",
"https://x.com/u",
"My title",
"My content",
"tag1 tag2",
)),
false,
"<a href=\"https://x.com/u\">Author</a> · My title",
&media,
);
assert!(report.contains("site: twitter"), "{report}");
assert!(report.contains("key: twitter:1"), "{report}");
assert!(report.contains("title: My title"), "{report}");
assert!(report.contains("content: My content"), "{report}");
assert!(report.contains("author: Author"), "{report}");
assert!(report.contains("author_url: https://x.com/u"), "{report}");
assert!(report.contains("tags: tag1 tag2"), "{report}");
assert!(report.contains("sensitive: false"), "{report}");
assert!(report.contains("media (2):"), "{report}");
assert!(
report.contains("1. image: https://cdn.example/1.jpg"),
"{report}"
);
assert!(
report.contains("2. video: https://cdn.example/2.mp4"),
"{report}"
);
}
#[test]
fn debug_report_without_render_data_and_no_media() {
let report = debug_report("u", "pixiv", "s", "t", "c", None, true, "p", &[]);
assert!(!report.contains("author:"), "{report}");
assert!(report.contains("sensitive: true"), "{report}");
assert!(report.contains("media (0):"), "{report}");
}
#[test]
fn debug_report_wraps_caption_in_blockquote() {
// The report is an HTML message: raw fields are escaped, pre-escaped
// render fields are embedded as-is, and the caption is wrapped in a
// <blockquote> so it shows exactly as it will render in the sent
// media caption (escaped text and links included).
let report = debug_report(
"https://x.com/u/status/1",
"twitter",
"https://x.com/u/status/1",
"A & B <C>",
"body & <more>",
Some((
"A &amp; B",
"https://x.com/u",
"A &amp; B &lt;C&gt;",
"body &amp; &lt;more&gt;",
"#a &amp; #b",
)),
false,
"<a href=\"https://x.com/u\">A &amp; B</a>: C &lt;D&gt; &amp; E",
&[],
);
// Raw fields escaped (they render back to the original text in HTML).
assert!(report.contains("title: A &amp; B &lt;C&gt;"), "{report}");
assert!(
report.contains("source_url: https://x.com/u/status/1"),
"{report}"
);
// Pre-escaped render fields embedded as-is.
assert!(report.contains("author: A &amp; B"), "{report}");
assert!(report.contains("tags: #a &amp; #b"), "{report}");
// Caption wrapped in a blockquote with its HTML preserved.
assert!(
report.contains(
"caption: <blockquote><a href=\"https://x.com/u\">A &amp; B</a>: C &lt;D&gt; &amp; E</blockquote>"
),
"{report}"
);
}
#[test]
fn debug_report_is_capped() {
// 200 media lines ≈ 8 KB, comfortably over the cap.
let media: Vec<Media> = (0..200)
.map(|i| Media::Illustration {
title: None,
url: format!("https://cdn.example/{i}.jpg"),
thumbnail_url: None,
fallback_url: None,
})
.collect();
let report = debug_report("u", "twitter", "s", "t", "c", None, false, "p", &media);
assert!(report.chars().count() <= MAX_DEBUG_REPORT_CHARS, "{report}");
assert!(report.ends_with('…'), "{report}");
}
}
+249
View File
@@ -0,0 +1,249 @@
//! Inline query handling with a keystroke debounce: only a query stable for
//! [`INLINE_DEBOUNCE`] triggers a fetch, and repeats are served by Telegram's
//! inline cache instead of re-fetching.
use super::log_key;
use std::collections::HashMap;
use std::sync::LazyLock;
use teloxide::RequestError;
use teloxide::prelude::*;
use teloxide::types::{
InlineQuery, InlineQueryResult, InlineQueryResultMpeg4Gif, InlineQueryResultPhoto,
InlineQueryResultVideo, ParseMode,
};
use x_media::media::Media;
/// Debounce window for inline queries: Telegram fires an inline query on
/// every keystroke, and each prefix of a pasted URL (e.g. `.../status/12`,
/// `.../status/123`, ...) already matches the site patterns. Without a
/// debounce every keystroke triggers a fetch (3 attempts!) of a half-typed
/// post id. Only answer once the query has been stable for this long.
const INLINE_DEBOUNCE: std::time::Duration = std::time::Duration::from_millis(800);
/// Last seen inline query per user and whether it was already answered.
/// Guards the debounce timer: a repeat of an answered query is served by
/// Telegram's inline cache (see `cache_time`), not by another fetch. Keyed by
/// user id — a single shared slot would let one user's typing burst (or a
/// different user's query) cancel another user's pending answer.
struct InlineDebounceState {
query: String,
answered: bool,
}
#[derive(Default)]
struct DebounceStates(HashMap<u64, InlineDebounceState>);
impl DebounceStates {
/// Records `query` as the user's newest query. Returns false when it is a
/// repeat whose answer already went out (Telegram's inline cache serves
/// it; re-fetching would only hit the source site again).
fn note(&mut self, user_id: u64, query: &str) -> bool {
if let Some(prev) = self.0.get(&user_id)
&& prev.query == query
&& prev.answered
{
return false;
}
self.0.insert(
user_id,
InlineDebounceState {
query: query.to_string(),
answered: false,
},
);
true
}
/// Claims the answer for the user's newest query; false when a newer query
/// superseded it or the answer was already claimed.
fn claim(&mut self, user_id: u64, query: &str) -> bool {
let Some(state) = self.0.get_mut(&user_id) else {
return false;
};
if state.query != query || state.answered {
return false;
}
state.answered = true;
true
}
/// Releases a claimed-but-unsent answer so a repeat can retry the fetch.
fn release(&mut self, user_id: u64, query: &str) {
if let Some(state) = self.0.get_mut(&user_id)
&& state.query == query
{
state.answered = false;
}
}
}
static INLINE_DEBOUNCE_STATE: LazyLock<parking_lot::Mutex<DebounceStates>> =
LazyLock::new(|| parking_lot::Mutex::new(DebounceStates::default()));
pub async fn inline_query_handler(bot: Bot, query: InlineQuery) -> Result<(), RequestError> {
if query.query.is_empty() {
return respond(());
}
// Only run a fetch for something that is actually a supported post URL.
if x_media::site::cache_key(&query.query).is_none() {
return respond(());
}
// Debounce: record the query and answer only after it has been stable for
// INLINE_DEBOUNCE (the timer below). An already-answered repeat of the
// same query is left to Telegram's inline cache instead of re-fetching.
let user_id = query.from.id.0;
if !INLINE_DEBOUNCE_STATE.lock().note(user_id, &query.query) {
return respond(());
}
let query_text = query.query.clone();
tokio::spawn(async move {
tokio::time::sleep(INLINE_DEBOUNCE).await;
// Only the user's last query of a typing burst survives: earlier
// timers see the query changed and give up without answering.
if !INLINE_DEBOUNCE_STATE.lock().claim(user_id, &query_text) {
return;
}
match answer_inline_query(bot, query).await {
Ok(true) => {}
// No results produced (or nothing to answer): let a repeat of the
// same query retry the fetch.
Ok(false) | Err(_) => INLINE_DEBOUNCE_STATE.lock().release(user_id, &query_text),
}
});
respond(())
}
/// Fetches the post behind an inline query and answers it. The caller has
/// already applied the debounce. Returns `true` when an answer was sent.
async fn answer_inline_query(bot: Bot, query: InlineQuery) -> Result<bool, RequestError> {
log::debug!(
"inline query: {} [key={}]",
query.query,
log_key(&query.query)
);
// No retries: the debounce plus a 1s/2s backoff would outlast the inline
// query the answer belongs to.
match x_media::site::fetch_once(&query.query).await {
Ok(Some(fetched)) => {
let mut results: Vec<InlineQueryResult> = Vec::new();
// Inline results have the same 1024-char caption limit as regular
// messages; truncate once here for all items, then apply the same
// long-post quoting as the send paths. `answer_inline_query` has no
// `AppContext` (the debounce spawns it), so the parsed config comes
// from the process-wide static, and the text is the *escaped*
// title/content the built-in caption embeds (the raw
// `Fetched.title`/`content` differ whenever the post contains
// `<`/`&`).
let caption = x_media::site::truncate_caption(&fetched.caption);
let text = fetched
.render_fields()
.map(|(_, _, title, content, _)| x_media::site::compose_text(title, content))
.unwrap_or_default();
let caption = crate::send::quote_long_caption(
&caption,
&text,
super::CONFIG.caption_quote_text_chars,
);
for (i, media) in fetched.media.iter().enumerate() {
let id = format!("{i}");
let Some(url) = url::Url::parse(media.url()).ok() else {
continue;
};
let thumbnail = media
.thumbnail_url()
.and_then(|t| url::Url::parse(t).ok())
.unwrap_or_else(|| url.clone());
let caption = caption.clone().into_owned();
let result = match media {
Media::Illustration { .. } => {
// Inline photo results have their own (smaller) size
// cap; use the reduced variant when one exists.
let photo_url = media
.smaller_url()
.and_then(|u| url::Url::parse(u).ok())
.unwrap_or_else(|| url.clone());
InlineQueryResult::Photo(
InlineQueryResultPhoto::new(id, photo_url, thumbnail)
.caption(caption)
.parse_mode(ParseMode::Html),
)
}
Media::Video { .. } => InlineQueryResult::Video(
InlineQueryResultVideo::new(
id,
url,
"video/mp4".parse().expect("valid mime"),
thumbnail,
fetched.title.clone(),
)
.caption(caption)
.parse_mode(ParseMode::Html),
),
Media::Animated { .. } => InlineQueryResult::Mpeg4Gif(
InlineQueryResultMpeg4Gif::new(id, url, thumbnail)
.caption(caption)
.parse_mode(ParseMode::Html),
),
};
results.push(result);
}
if !results.is_empty() {
// Explicit cache window: repeats of the same query within 5
// minutes are served by Telegram without hitting the bot.
bot.answer_inline_query(query.id, results)
.cache_time(300)
.await?;
return Ok(true);
}
}
Ok(None) => {}
Err(e) => log::error!("inline fetch {}: {e}", query.query),
}
Ok(false)
}
#[cfg(test)]
mod tests {
use super::DebounceStates;
const URL_A: &str = "https://x.com/a/status/1";
const URL_B: &str = "https://x.com/b/status/2";
#[test]
fn debounce_state_is_per_user() {
let mut states = DebounceStates::default();
// Two users query different links: both proceed, and neither timer
// cancels the other (a single shared slot dropped one of them).
assert!(states.note(1, URL_A));
assert!(states.note(2, URL_B));
assert!(states.claim(1, URL_A), "user 1's answer was cancelled");
assert!(states.claim(2, URL_B), "user 2's answer was cancelled");
}
#[test]
fn answered_query_is_suppressed_per_user_only() {
let mut states = DebounceStates::default();
assert!(states.note(1, URL_A));
assert!(states.claim(1, URL_A));
// A repeat of the answered query by the same user is left to
// Telegram's inline cache.
assert!(!states.note(1, URL_A));
// Another user pasting the same link still gets an answer.
assert!(states.note(2, URL_A));
assert!(states.claim(2, URL_A));
}
#[test]
fn newer_query_supersedes_and_failed_answer_is_released() {
let mut states = DebounceStates::default();
assert!(states.note(1, URL_A));
assert!(states.note(1, URL_B));
// The stale timer for the half-typed query gives up…
assert!(!states.claim(1, URL_A));
// …and the newest one answers.
assert!(states.claim(1, URL_B));
// No results → release so a repeat may retry the fetch.
states.release(1, URL_B);
assert!(states.claim(1, URL_B));
}
}
+274
View File
@@ -0,0 +1,274 @@
//! Update handlers and the per-URL media pipeline.
//!
//! Split into per-concern modules: [`commands`] (the `/`-command executor),
//! [`urls`] (URL extraction + the bounded worker pool + send dispatch),
//! [`inline`] (debounced inline queries), [`callback`] (edit-before-forward
//! buttons) and [`statics`] (the shared process-wide stores). This module
//! holds the message entry point and the helpers the others share.
mod callback;
mod commands;
mod inline;
mod statics;
mod urls;
pub use callback::callback_query_handler;
pub use commands::register_commands;
pub use inline::inline_query_handler;
pub use statics::{CHAT_STORE, CONFIG, LINK_CACHE, TASK_QUEUE};
pub use urls::{start_url_workers, stop_url_workers};
use crate::ctx::AppContext;
use crate::media_sender::MediaSender;
use commands::{Command, execute_command};
use teloxide::RequestError;
use teloxide::prelude::*;
use teloxide::types::{ChatId, ChatKind, Message, MessageId, ParseMode, ReplyParameters};
use teloxide::utils::command::BotCommands;
use urls::{URL_JOBS, extract_urls};
/// Reply to a message by id, keeping the reply decoration even if the
/// original was already deleted. Returns the reply's message id.
pub(crate) async fn reply(
sender: &dyn MediaSender,
chat_id: i64,
reply_to: MessageId,
text: impl Into<String>,
) -> Result<i64, RequestError> {
sender
.send_message(ChatId(chat_id), text.into(), Some(reply_to), None)
.await
}
/// Reply to a message by id with HTML parse mode (same reply decoration as
/// [`reply`]). Used by `/test`, whose report is an HTML message (the caption
/// is wrapped in a `<blockquote>` to show it exactly as it will render).
pub(crate) async fn reply_html(
bot: &Bot,
chat_id: i64,
reply_to: MessageId,
text: String,
) -> Result<i64, RequestError> {
// `<Bot as Requester>::` disambiguates from the MediaSender trait's
// same-named method (see media_sender.rs).
<Bot as Requester>::send_message(bot, ChatId(chat_id), text)
.parse_mode(ParseMode::Html)
.reply_parameters(ReplyParameters::new(reply_to).allow_sending_without_reply())
.await
.map(|message| message.id.0 as i64)
}
/// Log prefix tying the whole lifecycle of one link (fetch → send → cache →
/// forward) together: the normalized cache key (`twitter:123…`, `pixiv:123`,
/// `bsky:handle/rkey`, `bilibili:123…`) instead of the raw URL, so logs stay
/// short and do not echo full user-submitted URLs at info level.
pub fn log_key(url: &str) -> String {
x_media::site::cache_key(url).unwrap_or_else(|| "<unsupported>".to_string())
}
/// Edit-before-forward: a reply to the prompt swaps the caption of the first
/// forwarded message. Returns true when the message was consumed as an edit.
/// Body of [`message_handler`]'s edit branch, without teloxide update types so
/// it can be driven by tests.
async fn edit_message_handler(
ctx: &AppContext<'_>,
chat_id: i64,
reply_to_message_id: i64,
text: &str,
) -> bool {
let chat_data = ctx.chat_store.get(chat_id).await;
let Some(edit) = chat_data.edit_message.get(&reply_to_message_id) else {
return false;
};
let Some(first_forward_id) = edit.forward_message_ids.first() else {
return false;
};
let link = format!(
"<a href=\"{0}\">{1}</a>",
html_escape::encode_double_quoted_attribute(&edit.url),
html_escape::encode_text(text)
);
let new_text = if edit.template.is_empty() {
link
} else {
chat_data
.template
.get(&edit.template)
.map(|template| template.replace("[]", &link))
.unwrap_or(link)
};
match ctx
.sender
.edit_message_caption(
ChatId(chat_id),
MessageId(*first_forward_id as i32),
new_text,
)
.await
{
Ok(()) => log::info!(
"edit-before-forward: caption swapped on message {first_forward_id} for prompt {reply_to_message_id}"
),
Err(e) => log::error!("edit_message_caption failed: {e}"),
}
true
}
pub async fn message_handler(bot: Bot, message: Message) -> Result<(), RequestError> {
let is_private = matches!(message.chat.kind, ChatKind::Private(_));
let sender = message
.from
.as_ref()
.map(|from| from.full_name())
.unwrap_or_else(|| "unknown".to_string());
let text_preview = message
.text()
.map(|t| {
let end = t.floor_char_boundary(120.min(t.len()));
&t[..end]
})
.unwrap_or("<no text>");
// Per-request detail: debug only (message text is user data).
log::debug!(
"message from {sender} in {} (private={is_private}): {text_preview}",
message.chat.id
);
// URL/edit flows only run in private chats; commands run in any chat.
if is_private
&& let Some(reply) = message.reply_to_message()
&& let Some(text) = message.text()
&& edit_message_handler(
&AppContext::from_statics(&bot),
message.chat.id.0,
reply.id.0 as i64,
text,
)
.await
{
return respond(());
}
if let Some(text) = message.text()
&& let Ok(command) = Command::parse(text, "")
{
log::debug!("command from {}: {text_preview}", message.chat.id);
execute_command(&bot, &message, command).await?;
return respond(());
}
if is_private {
let urls = extract_urls(&message);
if !urls.is_empty() {
// Debug only, and echo the normalized keys instead of the raw URLs.
let keys: Vec<String> = urls.iter().map(|u| log_key(u)).collect();
log::debug!("extracted {} URL(s): {keys:?}", urls.len());
}
for url in urls {
// Clone out of the lock: the parking_lot guard is !Send and must
// not be held across the await below.
let Some(tx) = URL_JOBS.lock().clone() else {
log::warn!("url workers not started; dropping link");
break;
};
// A closed channel means the workers are stopping (shutdown):
// report the dropped link instead of losing it silently.
if tx.send((message.clone(), url)).await.is_err() {
log::warn!("url workers stopped; dropping link");
break;
}
}
}
respond(())
}
#[cfg(test)]
mod tests {
use super::*;
use crate::ctx::test_support::TestStores;
use crate::media_sender::test_support::{MockSender, Outcome};
use crate::state::EditMessage;
use teloxide::ApiError;
const PROMPT_ID: i64 = 7;
const FORWARDED_ID: i64 = 9;
fn api_error() -> RequestError {
RequestError::Api(ApiError::Unknown("Bad Request: message not found".into()))
}
/// Seeds a prompt record; `template` names the chat template used for it
/// (empty = none, the caption gets the bare link).
async fn seed_prompt(ctx: &AppContext<'_>, template: &str) {
ctx.chat_store
.update(1, |data| {
data.template
.insert("tpl".to_string(), "<b>[]</b>".to_string());
data.edit_message.insert(
PROMPT_ID,
EditMessage {
url: "https://x.com/u/status/1".into(),
chat_id: 1,
forward_message_ids: vec![FORWARDED_ID],
template: template.to_string(),
created_at: crate::db::unix_now(),
},
);
})
.await;
}
#[tokio::test]
async fn reply_to_a_prompt_swaps_the_caption_through_its_template() {
let sender = MockSender::scripted(vec![Outcome::EditOk], api_error);
let stores = TestStores::new();
let ctx = stores.ctx(&sender);
seed_prompt(&ctx, "tpl").await;
let consumed = edit_message_handler(&ctx, 1, PROMPT_ID, "new caption").await;
assert!(consumed, "a reply to the prompt must be consumed");
assert_eq!(
sender.captions(),
vec!["<b><a href=\"https://x.com/u/status/1\">new caption</a></b>"]
);
}
#[tokio::test]
async fn reply_text_and_url_are_escaped_into_the_caption() {
let sender = MockSender::scripted(vec![Outcome::EditOk], api_error);
let stores = TestStores::new();
let ctx = stores.ctx(&sender);
seed_prompt(&ctx, "").await;
edit_message_handler(&ctx, 1, PROMPT_ID, "<script>alert(1)</script>").await;
// No raw markup from user text may reach the HTML caption.
assert_eq!(
sender.captions(),
vec!["<a href=\"https://x.com/u/status/1\">&lt;script&gt;alert(1)&lt;/script&gt;</a>"]
);
}
#[tokio::test]
async fn a_failed_caption_swap_still_consumes_the_reply() {
let sender = MockSender::scripted(vec![Outcome::EditErr], api_error);
let stores = TestStores::new();
let ctx = stores.ctx(&sender);
seed_prompt(&ctx, "tpl").await;
// The edit failed (message deleted etc.); the reply must still be
// swallowed instead of being treated as a link to fetch.
assert!(edit_message_handler(&ctx, 1, PROMPT_ID, "new caption").await);
assert_eq!(sender.calls(), vec!["edit_message_caption"]);
}
#[tokio::test]
async fn reply_to_an_unrelated_message_is_not_consumed() {
let sender = MockSender::scripted(vec![], api_error);
let stores = TestStores::new();
let ctx = stores.ctx(&sender);
// No prompt record for that message id → the reply runs the normal
// (URL/command) path instead.
assert!(!edit_message_handler(&ctx, 1, PROMPT_ID, "hello").await);
assert!(sender.calls().is_empty());
}
}
+38
View File
@@ -0,0 +1,38 @@
//! Process-wide singletons shared by the handler modules: the one SQLite
//! pool (and the three stores built on it) plus the configuration.
use crate::config::Config;
use crate::db::{self};
use crate::link_cache::LinkCache;
use crate::queue::PersistentTaskQueue;
use crate::state::ChatStore;
use std::sync::{Arc, LazyLock};
/// One shared SQLite pool for the three stores (chat state, task queue, link
/// cache): a single pool bounds concurrent DB work on `data/task_queue.db`
/// instead of three independent pools competing for the same file. The schema
/// for all three tables is initialized once, here.
static DB: LazyLock<Arc<db::DbPool>> = LazyLock::new(|| {
let path = db_path();
db::open_store(&path.to_string_lossy()).expect("failed to open database")
});
/// DB file location: `$DATA_DIR/task_queue.db` (default `data`, relative to
/// the working directory — keeps the docker-compose `./data` mount and local
/// runs unchanged). The directory is created if missing: SQLite does not
/// create parent dirs, so the old hardcoded `data/task_queue.db` failed with
/// a confusing error when started from a directory without `data/`, and a
/// CWD-relative path is a footgun for systemd / cron deployments — `DATA_DIR`
/// lets them pin the state anywhere.
fn db_path() -> std::path::PathBuf {
let dir = std::env::var("DATA_DIR").unwrap_or_else(|_| "data".to_string());
let dir_path = std::path::Path::new(&dir);
std::fs::create_dir_all(dir_path).expect("failed to create data directory");
dir_path.join("task_queue.db")
}
pub static CHAT_STORE: LazyLock<ChatStore> = LazyLock::new(|| ChatStore::new(Arc::clone(&DB)));
pub static TASK_QUEUE: LazyLock<PersistentTaskQueue> =
LazyLock::new(|| PersistentTaskQueue::new(Arc::clone(&DB)));
pub static LINK_CACHE: LazyLock<LinkCache> = LazyLock::new(|| LinkCache::new(Arc::clone(&DB)));
pub static CONFIG: LazyLock<Config> = LazyLock::new(Config::load);
+698
View File
@@ -0,0 +1,698 @@
//! URL extraction and the per-URL media pipeline: bounded job channel +
//! worker pool, link-cache fast path, fetch, task build and send dispatch.
use super::{log_key, reply};
use crate::ctx::{AppContext, CONTEXT};
use crate::link_cache::{CachedMediaKind, CachedPost};
use crate::send::{self, MediaItemPayload, Task};
use crate::state::ChatData;
use std::collections::HashSet;
use std::sync::LazyLock;
use teloxide::types::{ChatAction, ChatId, Message, MessageEntityKind, MessageId};
use x_media::media::Media;
/// One URL job: the message + the extracted URL (the sender and stores come
/// from the shared [`AppContext`], assembled from statics inside the worker).
type UrlJob = (Message, String);
/// Bounded channel of URL jobs drained by [`start_url_workers`]. The bound
/// caps both queued memory and shutdown backlog; a full channel applies
/// backpressure to the per-chat handler instead of spawning unbounded tasks.
pub(crate) static URL_JOBS: LazyLock<
parking_lot::Mutex<Option<tokio::sync::mpsc::Sender<UrlJob>>>,
> = LazyLock::new(|| parking_lot::Mutex::new(None));
/// Set by main's shutdown sequence; workers stop pulling new jobs.
pub(crate) static URL_STOP: std::sync::atomic::AtomicBool =
std::sync::atomic::AtomicBool::new(false);
/// JoinHandles of the URL workers, awaited by [`stop_url_workers`].
static URL_WORKER_HANDLES: LazyLock<parking_lot::Mutex<Option<Vec<tokio::task::JoinHandle<()>>>>> =
LazyLock::new(|| parking_lot::Mutex::new(None));
/// Worker count draining URL jobs; keeps the old 8-permit concurrency cap
/// while bounding how many jobs can be queued at all.
const URL_WORKERS: usize = 8;
/// Starts the URL job workers (called once from main after the queue starts).
/// teloxide dispatches updates to a per-chat worker that handles them
/// sequentially, so a batch-forward of many messages would otherwise be
/// processed one at a time (fetch + send each, roughly a second per
/// message); the workers add throughput, and FIFO order preserves per-message
/// URL order.
pub async fn start_url_workers() {
let (tx, rx) = tokio::sync::mpsc::channel::<UrlJob>(256);
*URL_JOBS.lock() = Some(tx);
let rx = std::sync::Arc::new(tokio::sync::Mutex::new(rx));
let mut handles = Vec::with_capacity(URL_WORKERS);
for _ in 0..URL_WORKERS {
let rx = std::sync::Arc::clone(&rx);
handles.push(tokio::spawn(async move {
while !URL_STOP.load(std::sync::atomic::Ordering::Relaxed) {
let job = rx.lock().await.recv().await;
match job {
Some((message, url)) => {
url_media(
&CONTEXT,
message.chat.id.0,
message.id.0 as i64,
&url,
PostSend::FromChat,
)
.await
}
None => break,
}
}
}));
}
*URL_WORKER_HANDLES.lock() = Some(handles);
}
/// Stops the URL workers: sets the stop flag, drops the job channel (so
/// workers blocked in \`recv()\` wake with \`None\` and exit) and awaits the
/// worker tasks. Each worker finishes its in-flight job first; jobs still
/// queued in the channel are abandoned (the old implementation neither
/// drained them nor woke blocked workers — it only set a flag checked
/// between jobs).
pub async fn stop_url_workers() {
URL_STOP.store(true, std::sync::atomic::Ordering::Relaxed);
// Dropping the sender makes every worker's recv() return None.
*URL_JOBS.lock() = None;
// Take the handles first so the lock guard drops before the awaits.
let handles = URL_WORKER_HANDLES.lock().take();
if let Some(handles) = handles {
for handle in handles {
let _ = handle.await;
}
}
}
/// Extracts URL and text-link entities (text + caption), deduped in order.
pub fn extract_urls(message: &Message) -> Vec<String> {
let mut urls = Vec::new();
for entity in message.parse_entities().into_iter().flatten() {
match entity.kind() {
MessageEntityKind::Url => urls.push(entity.text().to_string()),
MessageEntityKind::TextLink { url } => urls.push(url.to_string()),
_ => {}
}
}
for entity in message.parse_caption_entities().into_iter().flatten() {
match entity.kind() {
MessageEntityKind::Url => urls.push(entity.text().to_string()),
MessageEntityKind::TextLink { url } => urls.push(url.to_string()),
_ => {}
}
}
let mut seen = HashSet::new();
// Dedup by the normalized post id so variant URLs of the same post
// (/status/1 vs /status/1/photo/1) are sent once; unsupported URLs fall
// back to exact-string dedup.
urls.retain(|url| seen.insert(x_media::site::cache_key(url).unwrap_or_else(|| url.clone())));
urls
}
/// For locally produced media (encoded ugoira MP4) the thumbnail URL is a
/// hotlink-protected remote URL Telegram may not fetch; let Telegram generate
/// its own thumbnail instead.
fn thumbnail_for(media: &Media) -> Option<String> {
let url = media.url();
if url.starts_with("http://") || url.starts_with("https://") {
// An empty thumbnail string (misskey video/gif files without a
// thumbnailUrl) must not reach Telegram; let it generate its own.
media
.thumbnail_url()
.map(str::to_string)
.filter(|t| !t.is_empty())
} else {
None
}
}
fn media_to_payload(media: &Media, sensitive: bool) -> MediaItemPayload {
let fallback_url = media.smaller_url().map(str::to_string);
match media {
// A gif inside a group becomes a video item; a lone gif takes the
// animation path (see url_media).
Media::Illustration { .. } => MediaItemPayload::Photo {
media: media.url().to_string(),
has_spoiler: sensitive,
fallback_url,
file_id: false,
},
Media::Video { .. } => MediaItemPayload::Video {
media: media.url().to_string(),
has_spoiler: sensitive,
thumbnail: thumbnail_for(media),
fallback_url,
file_id: false,
},
Media::Animated { .. } => MediaItemPayload::Video {
media: media.url().to_string(),
has_spoiler: sensitive,
thumbnail: thumbnail_for(media),
fallback_url,
file_id: false,
},
}
}
/// Sends a task and handles the outcome: post-send actions on success, retry
/// enqueue on retryable failure, reply + link-cache invalidation on
/// permanent failure (a stale cached file id must not repeat forever).
async fn dispatch_send(
ctx: &AppContext<'_>,
chat_id: i64,
reply_to: MessageId,
task: &Task,
url: &str,
) {
let result = match task {
Task::SendAnimation { .. } => send::send_animation(ctx, task).await,
Task::SendMediaSequence { .. } => send::send_media_sequence(ctx, task).await,
Task::ForwardMessages { .. } => unreachable!(),
};
match result {
Ok(message_ids) => {
log::info!(
"sent {} message(s) for [key={}]",
message_ids.len(),
log_key(url)
);
send::post_send_actions(ctx, task, message_ids).await;
send::settle_task(ctx, task, send::Settled::Sent).await;
}
Err(send::SendError::Retryable {
delay_seconds,
task,
}) => {
log::info!(
"send for [key={}] failed, queued for retry in {delay_seconds:.1}s",
log_key(url)
);
send::enqueue_retry(ctx.task_queue, *task, delay_seconds).await;
let _ = reply(
ctx.sender,
chat_id,
reply_to,
"Send failed. Task queued for retry.",
)
.await;
}
Err(send::SendError::Permanent {
message: err_message,
task,
}) => {
send::settle_task(ctx, &task, send::Settled::Failed).await;
log::error!("send for {url} failed permanently: {err_message}");
let _ = reply(
ctx.sender,
chat_id,
reply_to,
format!("Send failed: {err_message}"),
)
.await;
}
}
}
/// Whether a send also runs the chat's post-send actions. `/test` sends with
/// them suppressed so a test can never forward to the channel or open the
/// edit-before-forward prompt; a normal link uses whatever the chat is
/// configured with.
#[derive(Clone, Copy, PartialEq, Eq, Debug)]
pub(crate) enum PostSend {
/// Apply the chat's `forward_channel_id` / `edit_before_forward`.
FromChat,
/// Send only: no channel forward, no edit prompt.
Suppressed,
}
/// Builds the send task from ready-made items, sharing the payload shape
/// between the fresh-fetch, link-cache and `/test` paths.
#[allow(clippy::too_many_arguments)]
fn build_send_task(
chat_data: &ChatData,
chat_id: i64,
reply_to_message_id: i64,
source_url: String,
caption: String,
items: Vec<MediaItemPayload>,
cache_data: Option<CachedPost>,
post_send: PostSend,
) -> Task {
// Notification ids stay set in both modes: a queued retry that
// dead-letters should still tell the chat.
let (edit_before_forward, forward_channel_id) = match post_send {
PostSend::FromChat => (chat_data.edit_before_forward, chat_data.forward_channel_id),
PostSend::Suppressed => (false, None),
};
if items.len() == 1 && matches!(items[0], MediaItemPayload::Animation { .. }) {
Task::SendAnimation {
chat_id,
reply_to_message_id,
caption,
animation: items.into_iter().next().unwrap(),
source_url,
edit_before_forward,
forward_channel_id,
notify_chat_id: Some(chat_id),
notify_message_id: Some(reply_to_message_id),
cache_data,
}
} else {
Task::SendMediaSequence {
chat_id,
reply_to_message_id,
caption,
// Photos first so a mixed photo+video group starts with a photo
// (Telegram's sendMediaGroup rule); order within each kind is kept.
media_batches: send::chunk_media_items(send::photos_first(items)),
batch_index: 0,
sent_message_ids: vec![],
source_url,
edit_before_forward,
forward_channel_id,
notify_chat_id: Some(chat_id),
notify_message_id: Some(reply_to_message_id),
cache_data,
}
}
}
/// The per-URL pipeline: link cache → fetch → build → send → post-send.
///
/// `post_send` selects whether the chat's forward/edit settings apply: the URL
/// workers pass [`PostSend::FromChat`], the `/test` command
/// [`PostSend::Suppressed`]. Everything else (cache write, retry enqueue,
/// dead-letter notification) is identical.
pub(crate) async fn url_media(
ctx: &AppContext<'_>,
chat_id: i64,
reply_to_message_id: i64,
url: &str,
post_send: PostSend,
) {
let reply_to = MessageId(reply_to_message_id as i32);
if let Err(e) = ctx
.sender
.send_chat_action(ChatId(chat_id), ChatAction::Typing)
.await
{
log::error!("send_chat_action failed: {e}");
}
// Link cache: a post sent before is re-sent from Telegram file ids —
// no source-site request, no download, no upload. Keyed by the
// normalized post id so x.com / fxtwitter / /photo/N variants collide.
if let Some(key) = x_media::site::cache_key(url)
&& let Some(cached) = ctx.link_cache.get(&key, ctx.config.link_cache_ttl).await
{
log::debug!("link cache hit for {key}");
let chat_data = ctx.chat_store.get(chat_id).await;
// Cache keys are prefixed with the site id ("twitter:…"), matching
// the value a fresh fetch would read from Fetched::site_id.
let site = x_media::site::site_id_from_key(&key);
let format = chat_data
.message_format
.get(site)
.cloned()
.unwrap_or_default();
let caption = if format.is_empty() {
x_media::site::truncate_caption(&cached.caption)
} else {
x_media::site::caption_from_fields(
&format,
"",
&cached.url,
&cached.author,
&cached.author_url,
&cached.title,
&cached.content,
&cached.tags,
)
};
let items: Vec<MediaItemPayload> = cached
.media
.iter()
.map(|m| match m.kind {
CachedMediaKind::Photo => MediaItemPayload::Photo {
media: m.file_id.clone(),
has_spoiler: cached.sensitive,
fallback_url: None,
file_id: true,
},
CachedMediaKind::Video => MediaItemPayload::Video {
media: m.file_id.clone(),
has_spoiler: cached.sensitive,
thumbnail: None,
fallback_url: None,
file_id: true,
},
CachedMediaKind::Animation => MediaItemPayload::Animation {
media: m.file_id.clone(),
has_spoiler: cached.sensitive,
file_id: true,
},
})
.collect();
let task = build_send_task(
&chat_data,
chat_id,
reply_to_message_id,
cached.url.clone(),
caption,
items,
Some(cached),
post_send,
);
dispatch_send(ctx, chat_id, reply_to, &task, url).await;
return;
}
log::debug!("fetching {url} [key={}]", log_key(url));
match x_media::site::fetch(url).await {
// Unsupported links are ignored silently (Python parity).
Ok(None) => {
log::debug!("no site pattern matches {url}; ignoring");
}
// Retries exhausted: notify the user (Rust-only requirement 3).
Err(e) => {
log::error!("fetch {url}: {e}");
let _ = reply(
ctx.sender,
chat_id,
reply_to,
"Failed to fetch media from this link.",
)
.await;
}
Ok(Some(mut fetched)) => {
if fetched.media.is_empty() {
let _ = reply(
ctx.sender,
chat_id,
reply_to,
"No media found or media type is not supported.",
)
.await;
return;
}
let chat_data = ctx.chat_store.get(chat_id).await;
// Per-site caption format override (empty -> built-in caption).
let format = chat_data
.message_format
.get(fetched.site_name())
.cloned()
.unwrap_or_default();
let caption = fetched.caption_with(&format);
// Raw render data for the link cache; the send fills in the
// Telegram file ids and persists the entry.
let cache_data =
fetched
.render_fields()
.map(|(author, author_url, title, content, tags)| CachedPost {
url: fetched.source_url.clone(),
caption: fetched.caption.clone(),
title: title.to_string(),
content: content.to_string(),
author: author.to_string(),
author_url: author_url.to_string(),
tags: tags.to_string(),
sensitive: fetched.sensitive,
media: vec![],
});
let items: Vec<MediaItemPayload> = fetched
.media
.iter()
.map(|media| media_to_payload(media, fetched.sensitive))
.collect();
let task = build_send_task(
&chat_data,
chat_id,
reply_to_message_id,
fetched.source_url.clone(),
caption,
items,
cache_data,
post_send,
);
// Hand the keep-alive temp dir (ugoira / bsky remux MP4) to the
// retry registry: a queued retry runs after this function returns
// and the fetch's own TempDir is dropped, so without this the
// local file would be gone by the time the retry sends it.
if let Some(dir) = fetched.take_keep_alive() {
send::KEEP_ALIVE.lock().push(dir);
}
dispatch_send(ctx, chat_id, reply_to, &task, url).await;
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::ctx::test_support::TestStores;
use crate::link_cache::CachedMedia;
use crate::media_sender::test_support::{MockSender, Outcome};
use std::time::Duration;
use teloxide::{ApiError, RequestError};
fn permanent_error() -> RequestError {
RequestError::Api(ApiError::Unknown(
"Bad Request: message is not modified".into(),
))
}
fn cached_photo_entry() -> CachedPost {
CachedPost {
url: "https://x.com/u/status/1".into(),
caption: "cap".into(),
title: "t".into(),
content: "c".into(),
author: "a".into(),
author_url: "au".into(),
tags: "".into(),
sensitive: false,
media: vec![CachedMedia {
kind: CachedMediaKind::Photo,
file_id: "file-1".into(),
}],
}
}
#[tokio::test]
async fn cache_hit_sends_file_ids_and_invalidates_on_permanent_failure() {
let stores = TestStores::new();
let sender = MockSender::scripted(
vec![Outcome::GroupErr, Outcome::MessageErr],
permanent_error,
);
let ctx = stores.ctx(&sender);
stores
.link_cache()
.put("twitter:1", &cached_photo_entry())
.await;
url_media(&ctx, 1, 2, "https://x.com/u/status/1", PostSend::FromChat).await;
// The cached file id went out as a group send; the permanent failure
// then triggered the fire-and-forget reply (its mock error is fine).
assert_eq!(
sender.calls(),
vec!["send_chat_action", "send_media_group", "send_message"]
);
// The stale cache entry was invalidated so the next request re-fetches.
assert!(
stores
.link_cache()
.get("twitter:1", Duration::from_secs(3600))
.await
.is_none()
);
}
#[tokio::test]
async fn cache_hit_success_keeps_the_cache_entry() {
let stores = TestStores::new();
let sender = MockSender::scripted(vec![Outcome::GroupOk], permanent_error);
let ctx = stores.ctx(&sender);
stores
.link_cache()
.put("twitter:1", &cached_photo_entry())
.await;
url_media(&ctx, 1, 2, "https://x.com/u/status/1", PostSend::FromChat).await;
assert_eq!(sender.calls(), vec!["send_chat_action", "send_media_group"]);
// Success must not evict the entry.
assert!(
stores
.link_cache()
.get("twitter:1", Duration::from_secs(3600))
.await
.is_some()
);
}
/// The caption-quote threshold matches the post's text inside the caption,
/// so a long-text cache hit is quoted and a short-text one is not.
#[tokio::test]
async fn cache_hit_quotes_a_long_text_caption() {
let mut stores = TestStores::new();
stores.config_mut().caption_quote_text_chars = 3;
let prefix = "https://x.com/u/status/1\n<a href=\"au\">a</a>: ";
for (text, expected) in [
(
"abc",
format!("{prefix}<blockquote expandable>abc</blockquote>"),
),
("ab", format!("{prefix}ab")),
] {
let sender = MockSender::scripted(vec![Outcome::GroupOk], permanent_error);
let ctx = stores.ctx(&sender);
let mut entry = cached_photo_entry();
entry.caption = format!("{prefix}{text}");
entry.title = String::new();
entry.content = text.into();
stores.link_cache().put("twitter:1", &entry).await;
url_media(&ctx, 1, 2, "https://x.com/u/status/1", PostSend::FromChat).await;
assert_eq!(sender.captions(), vec![expected], "text {text:?}");
}
}
#[tokio::test]
async fn unsupported_url_is_ignored_silently() {
let stores = TestStores::new();
let sender = MockSender::scripted(vec![], permanent_error);
let ctx = stores.ctx(&sender);
// No cache key → the fetch dispatcher returns Ok(None) without any
// network; nothing is sent or replied.
url_media(
&ctx,
1,
2,
"https://example.com/not-a-post",
PostSend::FromChat,
)
.await;
assert_eq!(sender.calls(), vec!["send_chat_action"]);
}
// ── Send modes: the URL flow vs `/test` ─────────────────────────────
/// A chat that has both post-send actions configured.
async fn seed_post_send_settings(ctx: &AppContext<'_>) {
ctx.chat_store
.update(1, |data| {
data.forward_channel_id = Some(2);
data.edit_before_forward = true;
})
.await;
}
#[tokio::test]
async fn chat_settings_apply_to_the_normal_link_flow() {
let stores = TestStores::new();
let sender =
MockSender::scripted(vec![Outcome::GroupOk, Outcome::MessageOk], permanent_error);
let ctx = stores.ctx(&sender);
stores
.link_cache()
.put("twitter:1", &cached_photo_entry())
.await;
seed_post_send_settings(&ctx).await;
url_media(&ctx, 1, 2, "https://x.com/u/status/1", PostSend::FromChat).await;
// Media group, then the edit prompt (edit-before-forward wins over the
// channel forward, which only runs once the prompt is confirmed).
assert_eq!(
sender.calls(),
vec!["send_chat_action", "send_media_group", "send_message"]
);
}
#[tokio::test]
async fn test_mode_sends_the_media_without_forwarding_or_editing() {
let stores = TestStores::new();
// Only the group send is scripted: any forward (copy_messages) or edit
// prompt (send_message) would panic with "unexpected outcome".
let sender = MockSender::scripted(vec![Outcome::GroupOk], permanent_error);
let ctx = stores.ctx(&sender);
stores
.link_cache()
.put("twitter:1", &cached_photo_entry())
.await;
seed_post_send_settings(&ctx).await;
url_media(&ctx, 1, 2, "https://x.com/u/status/1", PostSend::Suppressed).await;
assert_eq!(sender.calls(), vec!["send_chat_action", "send_media_group"]);
// The send is otherwise ordinary: the post stays cached.
assert!(
stores
.link_cache()
.get("twitter:1", Duration::from_secs(3600))
.await
.is_some()
);
}
#[test]
fn send_mode_decides_whether_chat_actions_ride_along() {
let chat = ChatData {
forward_channel_id: Some(2),
edit_before_forward: true,
..ChatData::default()
};
let with_chat = build_send_task(
&chat,
1,
2,
"https://x.com/u/status/1".into(),
"cap".into(),
vec![],
None,
PostSend::FromChat,
);
let Task::SendMediaSequence {
edit_before_forward,
forward_channel_id,
..
} = with_chat
else {
panic!("expected a media sequence task");
};
assert!(edit_before_forward);
assert_eq!(forward_channel_id, Some(2));
let suppressed = build_send_task(
&chat,
1,
2,
"https://x.com/u/status/1".into(),
"cap".into(),
vec![],
None,
PostSend::Suppressed,
);
let Task::SendMediaSequence {
edit_before_forward,
forward_channel_id,
notify_chat_id,
..
} = suppressed
else {
panic!("expected a media sequence task");
};
assert!(!edit_before_forward, "`/test` must not open an edit prompt");
assert_eq!(forward_channel_id, None, "`/test` must not forward");
// Dead-letter notification still reaches the chat that asked.
assert_eq!(notify_chat_id, Some(1));
}
}
+84 -4
View File
@@ -5,7 +5,7 @@
//! key]. A repeated link is then answered entirely from local state — no
//! re-fetch of the source site, no re-upload — and no media file is stored
//! on disk (the file ids point at Telegram's servers). Entries expire after
//! [`Config::link_cache_ttl`]; a stale entry is dropped lazily on read and
//! `Config::link_cache_ttl`; a stale entry is dropped lazily on read and
//! by the periodic prune in `main`.
use crate::db::now_f64;
@@ -38,6 +38,10 @@ pub struct CachedPost {
/// override).
pub caption: String,
pub title: String,
/// The post's body text. Defaulted on read: entries written before the
/// title/content split carry it inside `title`.
#[serde(default)]
pub content: String,
pub author: String,
pub author_url: String,
pub tags: String,
@@ -78,9 +82,15 @@ impl LinkCache {
conn.execute("DELETE FROM link_cache WHERE url = ?1", params![key])?;
return Ok(None);
}
Ok(Some(serde_json::from_str::<CachedPost>(&payload).map_err(
|e| rusqlite::Error::ToSqlConversionFailure(Box::new(e)),
)?))
match serde_json::from_str::<CachedPost>(&payload) {
Ok(post) => Ok(Some(post)),
Err(e) => {
// Unreadable payload (e.g. an older schema): drop it
// instead of re-failing the parse on every later hit.
conn.execute("DELETE FROM link_cache WHERE url = ?1", params![key])?;
Err(rusqlite::Error::ToSqlConversionFailure(Box::new(e)))
}
}
})
.await;
match result {
@@ -176,6 +186,7 @@ mod tests {
url: "https://x.com/u/status/1".into(),
caption: "cap".into(),
title: "t".into(),
content: "c".into(),
author: "a".into(),
author_url: "au".into(),
tags: "".into(),
@@ -187,6 +198,49 @@ mod tests {
}
}
/// A payload written before the title/content split has no `content`
/// field. It must still read back — the cache deletes what it cannot
/// parse — with its text left where it was stored (`title`) and the
/// caption it replays untouched. No migration: a self-hosted cache entry
/// lives one TTL, and moving the text would only reshuffle `/set_format`
/// placeholders until it expires.
#[tokio::test]
async fn pre_split_entry_still_parses() {
let dir = tempfile::tempdir().unwrap();
let cache = LinkCache::new(
crate::db::open_store(dir.path().join("c.db").to_str().unwrap()).unwrap(),
);
let legacy = serde_json::json!({
"url": "https://x.com/u/status/1",
"caption": "https://x.com/u/status/1\n<a href=\"au\">a</a>: old text",
"title": "old text",
"author": "a",
"author_url": "au",
"tags": "",
"sensitive": false,
"media": [{"kind": "photo", "file_id": "AgAC..."}]
});
{
let conn = rusqlite::Connection::open(dir.path().join("c.db")).unwrap();
conn.execute(
"INSERT INTO link_cache (url, payload, created_at) VALUES (?1, ?2, ?3)",
params!["twitter:1", legacy.to_string(), now_f64()],
)
.unwrap();
}
let got = cache
.get("twitter:1", Duration::from_secs(3600))
.await
.expect("a pre-split payload must not be dropped");
assert_eq!(got.title, "old text");
assert_eq!(got.content, "");
assert_eq!(
got.caption,
"https://x.com/u/status/1\n<a href=\"au\">a</a>: old text"
);
}
#[tokio::test]
async fn put_get_roundtrip() {
let dir = tempfile::tempdir().unwrap();
@@ -228,6 +282,32 @@ mod tests {
);
}
#[tokio::test]
async fn unreadable_entry_is_dropped_on_read() {
// A payload from an older schema must not be re-parsed on every hit:
// the row is removed and the read reports a miss.
let dir = tempfile::tempdir().unwrap();
let db_path = dir.path().join("c.db");
let cache = LinkCache::new(crate::db::open_store(db_path.to_str().unwrap()).unwrap());
{
let conn = rusqlite::Connection::open(&db_path).unwrap();
conn.execute(
"INSERT INTO link_cache (url, payload, created_at) VALUES (?1, ?2, ?3)",
params!["twitter:1", "{not json", now_f64()],
)
.unwrap();
}
assert!(
cache
.get("twitter:1", Duration::from_secs(3600))
.await
.is_none()
);
// Dropped, not left behind for the next hit.
assert_eq!(cache.clear(None).await, 0, "corrupted row still present");
}
#[tokio::test]
async fn remove_and_prune() {
let dir = tempfile::tempdir().unwrap();
+14 -2
View File
@@ -8,14 +8,18 @@ use tokio::sync::watch;
use x_media::site;
mod config;
mod ctx;
mod db;
mod handlers;
mod link_cache;
mod media_sender;
mod photo;
mod queue;
mod rate_limit;
mod send;
mod state;
use ctx::CONTEXT;
use handlers::{CHAT_STORE, CONFIG, LINK_CACHE, TASK_QUEUE};
/// Docker `stop` / `compose down` delivers SIGTERM, which teloxide's ctrlc
@@ -59,9 +63,13 @@ async fn main() {
);
// Queue worker: handles typed tasks, dead-letters failed sends to the
// task's chat.
// task's chat. Both closures use the shared context (the queue requires
// 'static handlers, and the statics are process-wide anyway).
TASK_QUEUE
.start(send::handle_task, send::dead_letter_notify)
.start(
|payload| send::handle_task(&CONTEXT, payload),
|payload, message| send::dead_letter_notify(&CONTEXT, payload, message),
)
.await;
log::info!("task queue worker started");
@@ -106,6 +114,10 @@ async fn main() {
if pruned > 0 {
log::info!("link cache: pruned {pruned} expired entr(ies)");
}
let idle_limiters = crate::rate_limit::prune_idle();
if idle_limiters > 0 {
log::debug!("rate limiter: dropped {idle_limiters} idle bucket(s)");
}
for (chat_id, prompt_message_id) in removed {
// If the prompt was already deleted, this fails with a
// 400 "message to edit not found" — log and ignore.
+453
View File
@@ -0,0 +1,453 @@
//! Send abstraction: the message-sending surface [`send`](crate::send)
//! needs, so the send pipeline can be tested with a scripted mock instead of
//! a live teloxide `Bot`.
use std::future::Future;
use std::pin::Pin;
use teloxide::RequestError;
use teloxide::prelude::Requester;
use teloxide::prelude::*;
use teloxide::types::{
CallbackQueryId, ChatAction, ChatId, InlineKeyboardMarkup, InputFile, InputMedia, Message,
MessageId, ParseMode, ReplyParameters,
};
/// Boxed, `Send` future returned by a [`MediaSender`] method (`async fn` in
/// traits is not dyn-compatible).
type BoxFuture<'a, T> = Pin<Box<dyn Future<Output = T> + Send + 'a>>;
/// The message-sending surface the send pipeline uses. The production
/// implementation is teloxide's [`Bot`]; tests inject a scripted mock to
/// cover the fallback and classification logic without touching the
/// Telegram API.
pub trait MediaSender: Send + Sync {
/// Sends a media group, replying to `reply_to`.
fn send_media_group(
&self,
chat_id: ChatId,
reply_to: MessageId,
items: Vec<InputMedia>,
) -> BoxFuture<'_, Result<Vec<Message>, RequestError>>;
/// Sends a lone animation, replying to `reply_to`.
fn send_animation<'a>(
&'a self,
chat_id: ChatId,
reply_to: MessageId,
caption: &'a str,
spoiler: bool,
file: InputFile,
) -> BoxFuture<'a, Result<Message, RequestError>>;
/// Copies messages between chats (forward to channel).
fn copy_messages(
&self,
to: ChatId,
from: ChatId,
ids: Vec<MessageId>,
) -> BoxFuture<'_, Result<Vec<MessageId>, RequestError>>;
/// Sends a plain text message, optionally replying to `reply_to` and
/// attaching `reply_markup`. Returns the sent message's id: the bot only
/// ever needs that (the edit-before-forward prompt's record is keyed by
/// it), and returning the whole `Message` would force every test mock to
/// construct one.
fn send_message(
&self,
chat_id: ChatId,
text: String,
reply_to: Option<MessageId>,
reply_markup: Option<InlineKeyboardMarkup>,
) -> BoxFuture<'_, Result<i64, RequestError>>;
/// Answers a callback query, optionally with a toast `text` shown to the
/// user who pressed the button.
fn answer_callback_query(
&self,
id: CallbackQueryId,
text: Option<String>,
) -> BoxFuture<'_, Result<(), RequestError>>;
/// Rewrites a message's caption, always with HTML parse mode (every caller
/// in this bot renders escaped HTML: templates and edit-before-forward
/// links).
fn edit_message_caption(
&self,
chat_id: ChatId,
message_id: MessageId,
caption: String,
) -> BoxFuture<'_, Result<(), RequestError>>;
/// Deletes a message (the edit-before-forward prompt after a forward).
fn delete_message(
&self,
chat_id: ChatId,
message_id: MessageId,
) -> BoxFuture<'_, Result<(), RequestError>>;
/// Sets the chat's "typing / uploading …" indicator (cosmetic).
fn send_chat_action(
&self,
chat_id: ChatId,
action: ChatAction,
) -> BoxFuture<'_, Result<(), RequestError>>;
}
impl MediaSender for Bot {
fn send_media_group(
&self,
chat_id: ChatId,
reply_to: MessageId,
items: Vec<InputMedia>,
) -> BoxFuture<'_, Result<Vec<Message>, RequestError>> {
Box::pin(async move {
// Pace media sends per chat (one token per item) so bursts do not
// trip Telegram's flood control.
crate::rate_limit::limiter_for(chat_id.0)
.acquire(items.len() as f64)
.await;
// `<Bot as Requester>::` disambiguates from this trait's same-named
// method (teloxide's API lives in the `Requester` trait).
<Bot as Requester>::send_media_group(self, chat_id, items)
.reply_parameters(ReplyParameters::new(reply_to).allow_sending_without_reply())
.await
})
}
fn send_animation<'a>(
&'a self,
chat_id: ChatId,
reply_to: MessageId,
caption: &'a str,
spoiler: bool,
file: InputFile,
) -> BoxFuture<'a, Result<Message, RequestError>> {
Box::pin(async move {
crate::rate_limit::limiter_for(chat_id.0).acquire(1.0).await;
let mut request = <Bot as Requester>::send_animation(self, chat_id, file)
.caption(caption)
.parse_mode(ParseMode::Html)
.reply_parameters(ReplyParameters::new(reply_to).allow_sending_without_reply());
if spoiler {
request = request.has_spoiler(true);
}
request.await
})
}
fn copy_messages(
&self,
to: ChatId,
from: ChatId,
ids: Vec<MessageId>,
) -> BoxFuture<'_, Result<Vec<MessageId>, RequestError>> {
Box::pin(async move {
// Channel forwards are the burstiest path (batch copies); pace
// them per message against the channel's budget.
crate::rate_limit::limiter_for(to.0)
.acquire(ids.len() as f64)
.await;
<Bot as Requester>::copy_messages(self, to, from, ids).await
})
}
fn send_message(
&self,
chat_id: ChatId,
text: String,
reply_to: Option<MessageId>,
reply_markup: Option<InlineKeyboardMarkup>,
) -> BoxFuture<'_, Result<i64, RequestError>> {
Box::pin(async move {
let mut request = <Bot as Requester>::send_message(self, chat_id, text);
if let Some(reply_to) = reply_to {
request = request
.reply_parameters(ReplyParameters::new(reply_to).allow_sending_without_reply());
}
if let Some(markup) = reply_markup {
request = request.reply_markup(markup);
}
request.await.map(|message| message.id.0 as i64)
})
}
fn answer_callback_query(
&self,
id: CallbackQueryId,
text: Option<String>,
) -> BoxFuture<'_, Result<(), RequestError>> {
Box::pin(async move {
let mut request = <Bot as Requester>::answer_callback_query(self, id);
if let Some(text) = text {
request = request.text(text);
}
request.await.map(|_| ())
})
}
fn edit_message_caption(
&self,
chat_id: ChatId,
message_id: MessageId,
caption: String,
) -> BoxFuture<'_, Result<(), RequestError>> {
Box::pin(async move {
<Bot as Requester>::edit_message_caption(self, chat_id, message_id)
.caption(caption)
.parse_mode(ParseMode::Html)
.await
.map(|_| ())
})
}
fn delete_message(
&self,
chat_id: ChatId,
message_id: MessageId,
) -> BoxFuture<'_, Result<(), RequestError>> {
Box::pin(async move {
<Bot as Requester>::delete_message(self, chat_id, message_id)
.await
.map(|_| ())
})
}
fn send_chat_action(
&self,
chat_id: ChatId,
action: ChatAction,
) -> BoxFuture<'_, Result<(), RequestError>> {
Box::pin(async move {
// teloxide's `send_chat_action` returns `Result<True, _>` (its
// unit marker type); map the success to `()`.
<Bot as Requester>::send_chat_action(self, chat_id, action)
.await
.map(|_| ())
})
}
}
/// Test support: a scripted [`MediaSender`] mock (no Telegram API involved).
#[cfg(test)]
pub(crate) mod test_support {
use super::*;
use parking_lot::Mutex;
/// One scripted outcome, consumed front-to-back; the last entry repeats
/// for further calls of the same method kind.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub(crate) enum Outcome {
GroupOk,
GroupErr,
AnimationErr,
CopyOk,
CopyErr,
/// An error from `send_message` (replies are fire-and-forget, so an
/// error is fine for tests).
MessageErr,
/// A successful `send_message`, returning message id [`MockSender::SENT_ID`].
MessageOk,
EditOk,
EditErr,
}
/// Replays a script and records what was sent, so tests can assert the
/// user-visible text a path produced.
pub(crate) struct MockSender {
script: Mutex<Vec<Outcome>>,
cursor: Mutex<usize>,
calls: Mutex<Vec<&'static str>>,
messages: Mutex<Vec<String>>,
captions: Mutex<Vec<String>>,
answers: Mutex<Vec<Option<String>>>,
/// Builds the error every `*Err` outcome returns (RequestError is not
/// cloneable, so the factory recreates it per call).
error: Box<dyn Fn() -> RequestError + Send + Sync>,
}
impl MockSender {
/// The message id a successful `send_message` reports.
pub(crate) const SENT_ID: i64 = 1;
pub(crate) fn scripted(
script: Vec<Outcome>,
error: impl Fn() -> RequestError + Send + Sync + 'static,
) -> Self {
MockSender {
script: Mutex::new(script),
cursor: Mutex::new(0),
calls: Mutex::new(Vec::new()),
messages: Mutex::new(Vec::new()),
captions: Mutex::new(Vec::new()),
answers: Mutex::new(Vec::new()),
error: Box::new(error),
}
}
/// Method names in call order (e.g. `["send_media_group",
/// "send_media_group"]` proves the fallback re-sent).
pub(crate) fn calls(&self) -> Vec<&'static str> {
self.calls.lock().clone()
}
/// Texts of the plain messages sent, in order.
pub(crate) fn messages(&self) -> Vec<String> {
self.messages.lock().clone()
}
/// Captions passed to `edit_message_caption`, in order.
pub(crate) fn captions(&self) -> Vec<String> {
self.captions.lock().clone()
}
/// Toast texts of the answered callback queries, in order.
pub(crate) fn answers(&self) -> Vec<Option<String>> {
self.answers.lock().clone()
}
fn next(&self, kind: &'static str) -> Outcome {
self.calls.lock().push(kind);
let script = self.script.lock();
let mut cursor = self.cursor.lock();
if script.is_empty() {
panic!("mock script exhausted: {kind}");
}
let idx = (*cursor).min(script.len() - 1);
*cursor = idx + 1;
script[idx]
}
fn error(&self) -> RequestError {
(self.error)()
}
}
impl MediaSender for MockSender {
fn send_media_group(
&self,
_chat_id: ChatId,
_reply_to: MessageId,
items: Vec<InputMedia>,
) -> BoxFuture<'_, Result<Vec<Message>, RequestError>> {
Box::pin(async move {
// Record the captions exactly as Telegram receives them (only
// the first item of a group carries one), so tests can assert
// what a recipient sees.
self.captions
.lock()
.extend(items.iter().filter_map(|item| match item {
InputMedia::Photo(photo) => photo.caption.clone(),
InputMedia::Video(video) => video.caption.clone(),
InputMedia::Animation(animation) => animation.caption.clone(),
_ => None,
}));
match self.next("send_media_group") {
Outcome::GroupOk => Ok(Vec::new()),
Outcome::GroupErr => Err(self.error()),
other => panic!("unexpected outcome {other:?} for send_media_group"),
}
})
}
fn send_animation<'a>(
&'a self,
_chat_id: ChatId,
_reply_to: MessageId,
_caption: &'a str,
_spoiler: bool,
_file: InputFile,
) -> BoxFuture<'a, Result<Message, RequestError>> {
Box::pin(async move {
match self.next("send_animation") {
Outcome::AnimationErr => Err(self.error()),
other => panic!("unexpected outcome {other:?} for send_animation"),
}
})
}
fn copy_messages(
&self,
_to: ChatId,
_from: ChatId,
_ids: Vec<MessageId>,
) -> BoxFuture<'_, Result<Vec<MessageId>, RequestError>> {
Box::pin(async move {
match self.next("copy_messages") {
Outcome::CopyOk => Ok(vec![MessageId(1)]),
Outcome::CopyErr => Err(self.error()),
other => panic!("unexpected outcome {other:?} for copy_messages"),
}
})
}
fn send_message(
&self,
_chat_id: ChatId,
text: String,
_reply_to: Option<MessageId>,
_reply_markup: Option<InlineKeyboardMarkup>,
) -> BoxFuture<'_, Result<i64, RequestError>> {
Box::pin(async move {
self.messages.lock().push(text);
match self.next("send_message") {
Outcome::MessageOk => Ok(MockSender::SENT_ID),
Outcome::MessageErr => Err(self.error()),
other => panic!("unexpected outcome {other:?} for send_message"),
}
})
}
fn answer_callback_query(
&self,
_id: CallbackQueryId,
text: Option<String>,
) -> BoxFuture<'_, Result<(), RequestError>> {
// Always succeeds: the toast is cosmetic, so the script stays
// focused on the outcomes a test cares about.
Box::pin(async move {
self.calls.lock().push("answer_callback_query");
self.answers.lock().push(text);
Ok(())
})
}
fn edit_message_caption(
&self,
_chat_id: ChatId,
_message_id: MessageId,
caption: String,
) -> BoxFuture<'_, Result<(), RequestError>> {
Box::pin(async move {
self.captions.lock().push(caption);
match self.next("edit_message_caption") {
Outcome::EditOk => Ok(()),
Outcome::EditErr => Err(self.error()),
other => panic!("unexpected outcome {other:?} for edit_message_caption"),
}
})
}
fn delete_message(
&self,
_chat_id: ChatId,
_message_id: MessageId,
) -> BoxFuture<'_, Result<(), RequestError>> {
// Deletion is fire-and-forget in every caller; always succeeds.
Box::pin(async move {
self.calls.lock().push("delete_message");
Ok(())
})
}
fn send_chat_action(
&self,
_chat_id: ChatId,
_action: ChatAction,
) -> BoxFuture<'_, Result<(), RequestError>> {
Box::pin(async move {
self.calls.lock().push("send_chat_action");
Ok(())
})
}
}
}
+4 -2
View File
@@ -122,7 +122,7 @@ fn output_channels(color: png::ColorType) -> usize {
/// white; 16-bit per channel was already stripped to 8-bit at decode.
fn flatten_rgba_to_rgb(rgba: &[u8]) -> Vec<u8> {
let mut rgb = Vec::with_capacity(rgba.len() / 4 * 3);
for px in rgba.chunks_exact(4) {
for px in rgba.as_chunks::<4>().0 {
let a = px[3] as u32;
for v in &px[..3] {
// Over white: C = C*a/255 + 255*(1 - a/255).
@@ -178,7 +178,9 @@ fn encode_jpeg(pix: &PixBuf, w: u32, h: u32) -> Result<Vec<u8>, String> {
PixBuf::GrayAlpha(v) => {
// JPEG has no alpha: composite onto white, output as gray.
let gray: Vec<u8> = v
.chunks_exact(2)
.as_chunks::<2>()
.0
.iter()
.map(|px| {
let (g, a) = (px[0] as u32, px[1] as u32);
((g * a + 255 * (255 - a)) / 255).min(255) as u8
+56 -6
View File
@@ -41,7 +41,13 @@ type DeadLetter = dyn Fn(Value, String) -> BoxFuture<'static, ()> + Send + Sync;
pub struct PersistentTaskQueue {
pool: std::sync::Arc<crate::db::DbPool>,
/// Wakes the workers when a row becomes leasable. `notify_one` stores a
/// permit, so nothing else may share it: a waiter that is not a worker
/// (the sweep) can consume the permit and leave the due row pending until
/// the next enqueue.
notify: Arc<Notify>,
/// Wakes the lease-expiry sweep; `stop` is the only producer.
sweep_notify: Arc<Notify>,
stop: Arc<AtomicBool>,
worker: Mutex<Vec<JoinHandle<()>>>,
counter: AtomicU64,
@@ -88,6 +94,7 @@ impl PersistentTaskQueue {
Self {
pool,
notify: Arc::new(Notify::new()),
sweep_notify: Arc::new(Notify::new()),
stop: Arc::new(AtomicBool::new(false)),
worker: Mutex::new(Vec::new()),
counter: AtomicU64::new(0),
@@ -119,12 +126,12 @@ impl PersistentTaskQueue {
handles.push(tokio::spawn(worker.run_loop_supervised()));
}
// Periodic lease-expiry sweep: recovers rows a crashed/panicked
// worker left `in_progress` (the lock TTL bounds the wait). Woken by
// the same notify as the workers, so enqueue and stop interrupt the
// sleep; the first interval tick fires immediately (harmless extra
// recovery at startup).
// worker left `in_progress` (the lock TTL bounds the wait). Its own
// notify (not the workers'): sharing that one let this task consume a
// `notify_one` permit meant for a worker, which then slept through a
// due row until some later event. Only `stop` wakes it.
let sweep_pool = std::sync::Arc::clone(&self.pool);
let sweep_notify = Arc::clone(&self.notify);
let sweep_notify = Arc::clone(&self.sweep_notify);
let sweep_stop = Arc::clone(&self.stop);
handles.push(tokio::spawn(async move {
let mut interval = tokio::time::interval(Duration::from_secs(30));
@@ -150,6 +157,7 @@ impl PersistentTaskQueue {
pub async fn stop(&self) {
self.stop.store(true, Ordering::Relaxed);
self.notify.notify_waiters();
self.sweep_notify.notify_waiters();
let handles = std::mem::take(&mut *self.worker.lock());
for handle in handles {
let _ = handle.await;
@@ -307,6 +315,11 @@ impl QueueWorker {
}
}
/// Processes one leased row, keeping the lease alive while the handler
/// runs. Without the heartbeat a task longer than [`LOCK_TTL_SECONDS`]
/// (slow download, ugoira encode, rate-limited batch forward) would have
/// its lease expire mid-run; the expiry sweep would flip the row back to
/// `pending` and another worker would process it again — duplicate sends.
async fn process(&self, row: LeasedRow) {
let payload: Value = match serde_json::from_str(&row.payload) {
Ok(value) => value,
@@ -318,7 +331,8 @@ impl QueueWorker {
}
};
log::debug!("processing {} (attempt {})", row.id, row.attempts + 1);
match (self.handler)(payload).await {
let outcome = self.run_with_lease(&row.id, payload).await;
match outcome {
Ok(()) => {
log::debug!("task {} completed", row.id);
self.delete_row(&row.id).await;
@@ -351,6 +365,42 @@ impl QueueWorker {
}
}
/// Drives the handler to completion, refreshing the row's `locked_until`
/// every 30 s so the expiry sweep never re-leases a still-running task.
/// The heartbeat is part of this future, not a separate spawned task: if
/// the worker task dies (panic) the heartbeat dies with it and the sweep
/// recovers the row exactly as before.
async fn run_with_lease(&self, id: &str, payload: Value) -> Result<(), QueueError> {
let fut = (self.handler)(payload);
tokio::pin!(fut);
let mut interval = tokio::time::interval(Duration::from_secs(30));
// The first interval tick fires immediately; skip it (the lease was
// just set by lease_next).
interval.tick().await;
let id_owned = id.to_string();
loop {
tokio::select! {
result = &mut fut => return result,
_ = interval.tick() => {
let now = now_f64();
let id = id_owned.clone();
let result = self
.pool
.with_conn(move |conn| {
conn.execute(
"UPDATE tasks SET locked_until=?1 WHERE id=?2 AND status='in_progress'",
params![now + LOCK_TTL_SECONDS, id],
)
})
.await;
if let Err(e) = result {
log::error!("queue lease heartbeat failed: {e}");
}
}
}
}
}
async fn delete_row(&self, id: &str) {
let id = id.to_string();
let result = self
+184
View File
@@ -0,0 +1,184 @@
//! Per-chat token-bucket rate limiting.
//!
//! Telegram throttles bots that burst past a chat's message budget
//! (roughly 20 messages/min for channels/groups); today the bot absorbs
//! those 429s with queue retries. This limiter smooths the burst *before*
//! it reaches the API: media sends to a chat consume one token per
//! message, refilled at [`REFILL_PER_SEC`], so a batch forward paces itself
//! instead of tripping flood control. The queue retry stays as the safety
//! net for limits this bucket does not model (global per-bot limits etc.).
use parking_lot::Mutex;
use std::collections::HashMap;
use std::sync::Arc;
use std::sync::LazyLock;
use std::time::Duration;
/// Burst capacity: how many messages may be sent at once without waiting.
const CAPACITY: f64 = 20.0;
/// Sustained refill: ~20 messages per minute.
const REFILL_PER_SEC: f64 = 20.0 / 60.0;
struct State {
/// Current token balance; may go negative (debt from an acquire larger
/// than the capacity, repaid by subsequent refills).
tokens: f64,
last_refill: tokio::time::Instant,
}
/// A token bucket: at most `CAPACITY` tokens accumulate, refilled at
/// `REFILL_PER_SEC`. [`TokenBucket::acquire`] consumes `n` tokens, waiting
/// for the deficit (a single acquire may exceed the capacity and goes into
/// debt, which the refill repays).
pub struct TokenBucket {
capacity: f64,
refill_per_sec: f64,
state: Mutex<State>,
}
impl TokenBucket {
fn new(capacity: f64, refill_per_sec: f64) -> Self {
TokenBucket {
capacity,
refill_per_sec,
state: Mutex::new(State {
tokens: capacity,
last_refill: tokio::time::Instant::now(),
}),
}
}
/// Applies the elapsed refill to `state`. Shared by [`Self::acquire`] and
/// the idle check so the two cannot drift apart.
fn refill(&self, state: &mut State) {
let now = tokio::time::Instant::now();
let elapsed = now
.saturating_duration_since(state.last_refill)
.as_secs_f64();
// Refill up to the capacity; a debt (negative balance) is repaid
// before any surplus accumulates.
state.tokens = (state.tokens + elapsed * self.refill_per_sec).min(self.capacity);
state.last_refill = now;
}
/// Waits until `n` tokens are available, consuming them. The wait is
/// bounded: the deficit is committed as debt and repaid over time, so a
/// large acquire returns once its share of the refill budget has passed.
pub async fn acquire(&self, n: f64) {
// The parking_lot guard is confined to this block: only the plain
// `wait` duration crosses the await (a guard across an await point
// would make the future !Send).
let wait = {
let mut state = self.state.lock();
self.refill(&mut state);
if state.tokens >= n {
state.tokens -= n;
return;
}
// Commit the whole consumption now; the caller proceeds once the
// deficit's worth of refill time has passed.
let debt = n - state.tokens;
state.tokens = -debt;
debt / self.refill_per_sec
};
tokio::time::sleep(Duration::from_secs_f64(wait)).await;
}
/// True when the bucket has refilled to capacity: no debt outstanding, so
/// the chat has not sent anything recently.
fn is_idle(&self) -> bool {
let mut state = self.state.lock();
self.refill(&mut state);
state.tokens >= self.capacity
}
}
/// One limiter per chat, created on first use. Per-chat so one chat's burst
/// never throttles another.
static LIMITERS: LazyLock<Mutex<HashMap<i64, Arc<TokenBucket>>>> =
LazyLock::new(|| Mutex::new(HashMap::new()));
/// Returns the shared limiter for a chat, creating it on first use.
pub fn limiter_for(chat_id: i64) -> Arc<TokenBucket> {
LIMITERS
.lock()
.entry(chat_id)
.or_insert_with(|| Arc::new(TokenBucket::new(CAPACITY, REFILL_PER_SEC)))
.clone()
}
/// Drops limiters that are idle (refilled to capacity, so the chat has not
/// sent recently) and are not still held by an in-flight sender. The map
/// would otherwise keep one bucket per chat that ever sent media, forever.
/// Called from the periodic sweep; returns how many were dropped.
pub fn prune_idle() -> usize {
let mut limiters = LIMITERS.lock();
let before = limiters.len();
// Lock order map → bucket, the only order taken anywhere.
limiters.retain(|_, bucket| Arc::strong_count(bucket) > 1 || !bucket.is_idle());
before - limiters.len()
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn limiter_for_reuses_the_per_chat_bucket() {
let a = limiter_for(1);
let b = limiter_for(1);
let c = limiter_for(2);
assert!(Arc::ptr_eq(&a, &b), "same chat → same bucket");
assert!(!Arc::ptr_eq(&a, &c), "different chat → different bucket");
}
#[tokio::test(start_paused = true)]
async fn burst_is_consumed_instantly_then_refill_waits() {
let bucket = TokenBucket::new(3.0, 1.0);
// A burst within capacity passes without waiting.
bucket.acquire(3.0).await;
// The bucket is empty now; one token needs 1s of refill.
let start = tokio::time::Instant::now();
bucket.acquire(1.0).await;
assert!(
start.elapsed() >= Duration::from_secs(1),
"elapsed {:?}",
start.elapsed()
);
}
#[tokio::test(start_paused = true)]
async fn acquire_larger_than_capacity_waits_for_the_deficit() {
let bucket = TokenBucket::new(2.0, 1.0);
// 5 tokens with a capacity of 2: the 3-token deficit takes 3s.
let start = tokio::time::Instant::now();
bucket.acquire(5.0).await;
assert!(
start.elapsed() >= Duration::from_secs(3),
"elapsed {:?}",
start.elapsed()
);
}
#[tokio::test(start_paused = true)]
async fn prune_idle_drops_full_unheld_buckets_only() {
// Held by this task: kept even at full capacity, a sender has it.
let held = limiter_for(9_001);
assert!(held.is_idle(), "a fresh bucket is full");
// Only the map holds this one and it is full → dropped.
limiter_for(9_002);
// Mid-debt (an acquire larger than the capacity): kept.
{
let bucket = Arc::new(TokenBucket::new(CAPACITY, REFILL_PER_SEC));
bucket.state.lock().tokens = -1.0;
LIMITERS.lock().insert(9_003, bucket);
}
assert!(prune_idle() >= 1);
let limiters = LIMITERS.lock();
assert!(limiters.contains_key(&9_001), "held bucket pruned");
assert!(!limiters.contains_key(&9_002), "idle unheld bucket kept");
assert!(limiters.contains_key(&9_003), "indebted bucket pruned");
}
}
File diff suppressed because it is too large Load Diff
+128
View File
@@ -0,0 +1,128 @@
//! Payload → Telegram input types: `InputFile` selection (cached file id /
//! URL / local path), the per-kind `InputMedia` builders and the media-group
//! assembly with its caption rule.
use super::MediaItemPayload;
use teloxide::types::{
InputFile, InputMedia, InputMediaAnimation, InputMediaPhoto, InputMediaVideo, ParseMode,
};
fn parse_media_url(s: &str) -> Result<url::Url, String> {
url::Url::parse(s).map_err(|e| format!("invalid media URL: {e}"))
}
pub(super) fn item_url(item: &MediaItemPayload) -> &str {
match item {
MediaItemPayload::Photo { media, .. }
| MediaItemPayload::Video { media, .. }
| MediaItemPayload::Animation { media, .. } => media,
}
}
/// Remote http(s) URLs are handed to Telegram to fetch; everything else
/// (e.g. a locally encoded ugoira MP4) is uploaded directly.
pub(super) fn input_file_for(media: &str) -> Result<InputFile, String> {
if media.starts_with("http://") || media.starts_with("https://") {
Ok(InputFile::url(parse_media_url(media)?))
} else if !std::path::Path::new(media).exists() {
// A retried task may reference a temp file the original send's
// TempDir already cleaned up; fail fast and permanent instead of
// burning retries on a file that can never come back.
Err(format!("local media file missing: {media}"))
} else {
Ok(InputFile::file(media))
}
}
impl MediaItemPayload {
/// The input for a send: a cached file id goes out as `InputFile::file_id`
/// (no fetch, no upload), URLs go to Telegram, anything else is a local
/// path (transient upload fallback).
fn input_file(&self) -> Result<InputFile, String> {
match self {
MediaItemPayload::Photo {
media,
file_id: true,
..
}
| MediaItemPayload::Video {
media,
file_id: true,
..
}
| MediaItemPayload::Animation {
media,
file_id: true,
..
} => Ok(InputFile::file_id(media.clone().into())),
_ => input_file_for(item_url(self)),
}
}
}
pub(super) fn photo_media(file: InputFile, caption: Option<&str>, spoiler: bool) -> InputMedia {
let mut photo = InputMediaPhoto::new(file).parse_mode(ParseMode::Html);
if let Some(caption) = caption {
photo = photo.caption(caption);
}
if spoiler {
photo = photo.spoiler();
}
InputMedia::Photo(photo)
}
pub(super) fn video_media(file: InputFile, caption: Option<&str>, spoiler: bool) -> InputMedia {
let mut video = InputMediaVideo::new(file).parse_mode(ParseMode::Html);
if let Some(caption) = caption {
video = video.caption(caption);
}
if spoiler {
video = video.spoiler();
}
InputMedia::Video(video)
}
pub(super) fn animation_media(file: InputFile, caption: Option<&str>, spoiler: bool) -> InputMedia {
let mut animation = InputMediaAnimation::new(file).parse_mode(ParseMode::Html);
if let Some(caption) = caption {
animation = animation.caption(caption);
}
if spoiler {
animation = animation.spoiler();
}
InputMedia::Animation(animation)
}
/// Builds a media group from payloads; only the first item of the batch gets
/// the caption (Telegram rejects captions on later items).
pub(super) fn build_media_group(
batch: &[MediaItemPayload],
caption: Option<&str>,
) -> Result<Vec<InputMedia>, String> {
batch
.iter()
.enumerate()
.map(|(i, item)| {
let item_caption = if i == 0 { caption } else { None };
Ok(match item {
MediaItemPayload::Photo { has_spoiler, .. } => {
photo_media(item.input_file()?, item_caption, *has_spoiler)
}
MediaItemPayload::Video {
has_spoiler,
thumbnail,
..
} => {
let mut video = video_media(item.input_file()?, item_caption, *has_spoiler);
if let (Some(thumb), InputMedia::Video(v)) = (thumbnail, &mut video) {
*v = v.clone().thumbnail(input_file_for(thumb)?);
}
video
}
MediaItemPayload::Animation { has_spoiler, .. } => {
animation_media(item.input_file()?, item_caption, *has_spoiler)
}
})
})
.collect()
}
File diff suppressed because it is too large Load Diff
+366
View File
@@ -0,0 +1,366 @@
//! Everything around a send: the link-cache write that follows one, the
//! keep-alive registry for locally produced media, task settlement, the
//! post-send actions (edit prompt / channel forward) and the queue entry
//! points.
use super::{SendError, Task, forward_messages, send_animation, send_media_sequence};
use crate::ctx::AppContext;
use crate::db::{now_f64, unix_now};
use crate::handlers::log_key;
use crate::link_cache::{CachedMedia, CachedMediaKind, LinkCache};
use crate::media_sender::MediaSender;
use crate::queue::{PersistentTaskQueue, QueueError};
use crate::state::EditMessage;
use std::collections::HashMap;
use std::sync::LazyLock;
use teloxide::types::{ChatId, InlineKeyboardButton, InlineKeyboardMarkup, Message, MessageId};
/// Persists a successful send under the post's cache key. Only runs for a
/// fresh (non-resumed) task that carried raw cache data with no file ids yet.
pub(super) async fn cache_sent_task(ctx: &AppContext<'_>, task: &Task, media: Vec<CachedMedia>) {
let Some(cache_data) = task.cache_data() else {
return;
};
if !cache_data.media.is_empty() || media.is_empty() {
return;
}
let mut post = cache_data.clone();
post.media = media;
if let Some(key) = x_media::site::cache_key(&post.url) {
ctx.link_cache.put(&key, &post).await;
log::debug!("cached send for [key={}]", log_key(&post.url));
}
}
/// Persists a lone animation send under the post's cache key.
pub(super) async fn cache_animation_send(ctx: &AppContext<'_>, task: &Task, message: &Message) {
if let Some(file_id) = message.animation().map(|a| a.file.id.to_string()) {
cache_sent_task(
ctx,
task,
vec![CachedMedia {
kind: CachedMediaKind::Animation,
file_id,
}],
)
.await;
}
}
/// How a task ended. The two states differ only in whether a link-cache entry
/// may still be holding the (now unusable) media.
pub(crate) enum Settled {
Sent,
Failed,
}
/// Every path that ends a task's life — sent, permanently failed, or
/// dead-lettered after the last retry — funnels through here, so the cleanup a
/// settled task owes cannot be forgotten by a new path: release the keep-alive
/// temp media (retryable tasks keep it, they will be resent) and drop the
/// link-cache entry that a failed send's stale file ids would keep poisoning.
pub(crate) async fn settle_task(ctx: &AppContext<'_>, task: &Task, outcome: Settled) {
if matches!(outcome, Settled::Failed) {
invalidate_cache(ctx.link_cache, task).await;
}
release_keep_alive(task);
}
/// A cached Telegram file id failed permanently (stale/expired); drop the
/// cache entry so the next request re-fetches instead of repeating it.
async fn invalidate_cache(cache: &LinkCache, task: &Task) {
if task.is_cached_send()
&& let Some(url) = task.source_url()
&& let Some(key) = x_media::site::cache_key(url)
{
log::debug!("removing stale link cache entry for [key={}]", log_key(url));
cache.remove(&key).await;
}
}
/// Locally produced media files (ugoira MP4, bsky remux MP4) whose temp dirs
/// must stay alive while their task may be retried by the queue. The fetch
/// pipeline hands ownership here via
/// [`x_media::site::Fetched::take_keep_alive`] before that
/// [`x_media::site::Fetched`] is dropped; a queued retry runs after that drop,
/// so without this the local file would be gone by the time the retry sends
/// it. Entries are removed when the task settles (see [`release_keep_alive`]).
pub(crate) static KEEP_ALIVE: LazyLock<parking_lot::Mutex<Vec<tempfile::TempDir>>> =
LazyLock::new(|| parking_lot::Mutex::new(Vec::new()));
/// Drops the keep-alive temp dirs holding media referenced by `task` (matched
/// by path prefix). Called once a task settles — sent or permanently failed —
/// so retry-only temp files do not leak; retryable tasks keep them alive.
pub(crate) fn release_keep_alive(task: &Task) {
let paths = task.local_media_paths();
if paths.is_empty() {
return;
}
let mut alive = KEEP_ALIVE.lock();
alive.retain(|dir| {
let dir_path = dir.path();
!paths.iter().any(|p| p.starts_with(dir_path))
});
}
/// One button per template name (column layout), then the confirm button.
/// Sorted by name: the templates live in a `HashMap`, so an unsorted walk
/// would reshuffle the buttons between prompts.
pub(super) fn build_edit_markup(templates: &HashMap<String, String>) -> InlineKeyboardMarkup {
let mut names: Vec<&String> = templates.keys().collect();
names.sort();
let mut rows = Vec::with_capacity(names.len() + 1);
for name in names {
rows.push(vec![InlineKeyboardButton::callback(
name.clone(),
format!("template|{name}"),
)]);
}
rows.push(vec![InlineKeyboardButton::callback(
"↩️ Confirm",
"forward",
)]);
InlineKeyboardMarkup::new(rows)
}
/// Notifies a chat about a dead-lettered task (skips when `notify_chat_id` is
/// absent).
pub(super) async fn notify_failure(
sender: &dyn MediaSender,
chat_id: Option<i64>,
message_id: Option<i64>,
message: &str,
) {
let Some(chat_id) = chat_id else { return };
let reply_to = message_id.map(|id| MessageId(id as i32));
if let Err(e) = sender
.send_message(ChatId(chat_id), message.to_string(), reply_to, None)
.await
{
log::error!("failed to notify about failed task: {e}");
}
}
/// After a successful send: either open the edit-before-forward prompt or
/// forward to the configured channel (with retry/queue handling).
pub(crate) async fn post_send_actions(ctx: &AppContext<'_>, task: &Task, message_ids: Vec<i64>) {
let (
chat_id,
reply_to,
source_url,
edit_before_forward,
forward_channel_id,
notify_chat_id,
notify_message_id,
) = match task {
Task::SendMediaSequence {
chat_id,
reply_to_message_id,
source_url,
edit_before_forward,
forward_channel_id,
notify_chat_id,
notify_message_id,
..
}
| Task::SendAnimation {
chat_id,
reply_to_message_id,
source_url,
edit_before_forward,
forward_channel_id,
notify_chat_id,
notify_message_id,
..
} => (
*chat_id,
*reply_to_message_id,
source_url.clone(),
*edit_before_forward,
*forward_channel_id,
*notify_chat_id,
*notify_message_id,
),
Task::ForwardMessages { .. } => return,
};
if edit_before_forward {
let keyboard = build_edit_markup(&ctx.chat_store.get(chat_id).await.template);
let prompt = ctx
.sender
.send_message(
ChatId(chat_id),
"Reply to edit message.".to_string(),
Some(MessageId(reply_to as i32)),
Some(keyboard),
)
.await;
match prompt {
Ok(prompt_id) => {
log::info!(
"edit-before-forward prompt {prompt_id} opened for {} message(s)",
message_ids.len()
);
let source_url = source_url.clone();
ctx.chat_store
.update(chat_id, move |data| {
data.edit_message.insert(
prompt_id,
EditMessage {
url: source_url,
chat_id,
forward_message_ids: message_ids,
template: String::new(),
created_at: unix_now(),
},
);
})
.await;
}
Err(e) => log::error!("failed to send edit prompt: {e}"),
}
return;
}
if let Some(channel_id) = forward_channel_id {
log::info!(
"forwarding {} message(s) to channel {channel_id}",
message_ids.len()
);
let forward_task = Task::ForwardMessages {
from_chat_id: chat_id,
to_chat_id: channel_id,
message_ids,
notify_chat_id,
notify_message_id,
};
match forward_messages(ctx, &forward_task).await {
Ok(()) => {}
Err(SendError::Retryable {
delay_seconds,
task,
}) => {
enqueue_retry(ctx.task_queue, *task, delay_seconds).await;
}
Err(SendError::Permanent { message, .. }) => {
notify_failure(
ctx.sender,
notify_chat_id,
notify_message_id,
&format!("Task failed after retries: {message}"),
)
.await;
}
}
}
}
/// Enqueues a task for a later attempt (retry / forward resume). When the
/// enqueue itself fails the task can never be sent again, so its keep-alive
/// temp media is released instead of leaking until process exit.
pub(crate) async fn enqueue_retry(queue: &PersistentTaskQueue, task: Task, delay_seconds: f64) {
let payload = serde_json::to_value(&task).expect("task serializes");
let run_after = now_f64() + delay_seconds;
if let Err(e) = queue.enqueue(payload, run_after).await {
log::error!("failed to enqueue retry: {e}");
release_keep_alive(&task);
}
}
/// Queue entry point: parses the stored task and dispatches.
pub(crate) async fn handle_task(
ctx: &AppContext<'_>,
payload: serde_json::Value,
) -> Result<(), QueueError> {
let task: Task = match serde_json::from_value(payload.clone()) {
Ok(task) => task,
Err(e) => {
return Err(QueueError::Permanent {
message: format!("invalid task payload: {e}"),
payload,
});
}
};
match task {
Task::SendMediaSequence { .. } | Task::SendAnimation { .. } => {
let message_ids = match send_media_or_animation(ctx, &task).await {
Ok(ids) => ids,
Err(SendError::Retryable {
delay_seconds,
task,
}) => {
return Err(QueueError::Retryable {
delay_seconds,
payload: serde_json::to_value(task).expect("task serializes"),
});
}
Err(SendError::Permanent { message, task }) => {
settle_task(ctx, &task, Settled::Failed).await;
return Err(QueueError::Permanent {
message,
payload: serde_json::to_value(task).expect("task serializes"),
});
}
};
// A task only reaches the queue after a failed send, so this
// successful run is the first time post_send_actions can fire —
// the fresh attempt failed before it ever got here. Run it
// unconditionally: `post_send_actions` executes once, after the
// whole sequence (every batch) completed, so the channel forward
// and the edit-before-forward prompt must not be lost just
// because the send needed a retry.
post_send_actions(ctx, &task, message_ids).await;
settle_task(ctx, &task, Settled::Sent).await;
Ok(())
}
Task::ForwardMessages { .. } => match forward_messages(ctx, &task).await {
Ok(()) => Ok(()),
Err(SendError::Retryable {
delay_seconds,
task,
}) => Err(QueueError::Retryable {
delay_seconds,
payload: serde_json::to_value(task).expect("task serializes"),
}),
Err(SendError::Permanent { message, task }) => {
settle_task(ctx, &task, Settled::Failed).await;
Err(QueueError::Permanent {
message,
payload: serde_json::to_value(task).expect("task serializes"),
})
}
},
}
}
async fn send_media_or_animation(ctx: &AppContext<'_>, task: &Task) -> Result<Vec<i64>, SendError> {
match task {
Task::SendMediaSequence { .. } => send_media_sequence(ctx, task).await,
Task::SendAnimation { .. } => send_animation(ctx, task).await,
Task::ForwardMessages { .. } => unreachable!(),
}
}
/// Dead-letter callback wired to the queue in main: settles the task and
/// notifies its chat.
pub(crate) async fn dead_letter_notify(
ctx: &AppContext<'_>,
payload: serde_json::Value,
message: String,
) {
// A dead-lettered task never runs again, and the queue dead-letters retry
// exhaustion itself (the handler is not called again), so this is the only
// place that sees the final payload.
if let Ok(task) = serde_json::from_value::<Task>(payload.clone()) {
settle_task(ctx, &task, Settled::Failed).await;
}
let notify_chat_id = payload.get("notify_chat_id").and_then(|v| v.as_i64());
let notify_message_id = payload.get("notify_message_id").and_then(|v| v.as_i64());
notify_failure(
ctx.sender,
notify_chat_id,
notify_message_id,
&format!("Task failed after retries: {message}"),
)
.await;
}
+345
View File
@@ -0,0 +1,345 @@
//! Download-and-reupload fallback: when Telegram cannot fetch a media URL
//! itself (hotlink protection), the bot downloads the file, shrinks photos
//! that exceed Telegram's limits and uploads the batch via multipart.
use super::input_media::{animation_media, input_file_for, item_url, photo_media, video_media};
use super::{MediaItemPayload, SendError, Task, classify_to_send_error, retry_delay_seconds};
use crate::media_sender::MediaSender;
use crate::photo::{self, MAX_UPLOAD_BYTES, PhotoPrep};
use teloxide::prelude::*;
use teloxide::types::{ChatId, InputFile, InputMedia, MessageId};
use tempfile::NamedTempFile;
use x_media::site::FetchError;
/// Infers a file extension from magic bytes so Telegram detects the mime type
/// on multipart uploads.
pub(super) fn sniff_ext(bytes: &[u8]) -> &'static str {
if bytes.starts_with(&[0xFF, 0xD8]) {
"jpg"
} else if bytes.starts_with(b"\x89PNG") {
"png"
} else if bytes.starts_with(b"RIFF") && bytes.len() >= 12 && &bytes[8..12] == b"WEBP" {
"webp"
} else if bytes.starts_with(b"GIF8") {
"gif"
} else if bytes.len() >= 12 && &bytes[4..8] == b"ftyp" {
"mp4"
} else {
"bin"
}
}
pub(super) enum FallbackError {
Retryable {
delay_seconds: f64,
},
Permanent {
message: String,
},
/// The downloaded file exceeds the upload cap; the caller falls back to
/// the item's smaller URL.
MediaTooLarge,
}
/// Brings a downloaded photo within Telegram's limits via the pure-Rust
/// chain in [`crate::photo`] (no ffmpeg): dimension cap / upload cap
/// exceeded photos are decoded, downscaled with Lanczos3, PNG bit depth
/// reduced (>24-bit → 24-bit RGB, ≤24-bit untouched) and transcoded to JPEG
/// only if still too big. Anything that cannot be fixed falls back to the
/// item's smaller URL.
///
/// Downloads one media item to a temp file (deleted on drop), returning the
/// file plus the downloaded bytes (photos keep the bytes for
/// [`photo::prepare_photo`] — re-reading the file would double the I/O).
/// Network errors are retryable; size over the upload cap and other download
/// errors are not.
async fn download_to_temp(
item: &MediaItemPayload,
) -> Result<(NamedTempFile, bytes::Bytes), FallbackError> {
let media_url = match item {
MediaItemPayload::Photo { media, .. }
| MediaItemPayload::Video { media, .. }
| MediaItemPayload::Animation { media, .. } => media,
};
// Photos are downloaded even over the upload cap so `prepare_photo` can
// downscale / transcode them (cap = decode budget); videos/animations
// abort as soon as the upload cap is crossed mid-stream.
let limit = if matches!(item, MediaItemPayload::Photo { .. }) {
photo::MAX_DECODE_BYTES
} else {
MAX_UPLOAD_BYTES + 1
};
let bytes = match x_media::site::download_media_limited(media_url, limit).await {
Ok(bytes) => bytes,
Err(FetchError::Http(_)) => {
return Err(FallbackError::Retryable {
delay_seconds: retry_delay_seconds(0),
});
}
Err(FetchError::TooLarge) => {
return Err(FallbackError::MediaTooLarge);
}
Err(e) => {
return Err(FallbackError::Permanent {
message: format!("download failed: {e}"),
});
}
};
let ext = sniff_ext(&bytes);
let mut file = tempfile::Builder::new()
.suffix(&format!(".{ext}"))
.tempfile()
.map_err(|e| FallbackError::Permanent {
message: format!("temp file failed: {e}"),
})?;
use std::io::Write;
file.as_file_mut()
.write_all(&bytes)
.map_err(|e| FallbackError::Permanent {
message: format!("temp file write failed: {e}"),
})?;
Ok((file, bytes))
}
/// Builds the media group item from an uploaded file.
fn media_from_file(
item: &MediaItemPayload,
path: std::path::PathBuf,
caption: Option<&str>,
thumbnail: Option<&str>,
) -> Result<InputMedia, String> {
let mut media = match item {
MediaItemPayload::Photo { has_spoiler, .. } => {
photo_media(InputFile::file(path), caption, *has_spoiler)
}
MediaItemPayload::Video { has_spoiler, .. } => {
video_media(InputFile::file(path), caption, *has_spoiler)
}
MediaItemPayload::Animation { has_spoiler, .. } => {
animation_media(InputFile::file(path), caption, *has_spoiler)
}
};
if let (Some(thumb), InputMedia::Video(v)) = (thumbnail, &mut media) {
*v = v.clone().thumbnail(input_file_for(thumb)?);
}
Ok(media)
}
/// Builds the media group item from a (smaller) URL.
fn media_from_url(
item: &MediaItemPayload,
url: &str,
caption: Option<&str>,
thumbnail: Option<&str>,
) -> Result<InputMedia, String> {
let mut media = match item {
MediaItemPayload::Photo { has_spoiler, .. } => {
photo_media(input_file_for(url)?, caption, *has_spoiler)
}
MediaItemPayload::Video { has_spoiler, .. } => {
video_media(input_file_for(url)?, caption, *has_spoiler)
}
MediaItemPayload::Animation { has_spoiler, .. } => {
animation_media(input_file_for(url)?, caption, *has_spoiler)
}
};
if let (Some(thumb), InputMedia::Video(v)) = (thumbnail, &mut media) {
*v = v.clone().thumbnail(input_file_for(thumb)?);
}
Ok(media)
}
/// One item prepared for the upload fallback: the ready-to-send media plus
/// the temp file that must stay on disk until the group request completes.
pub(super) struct PreparedItem {
/// Original position in the batch (concurrent prep completes out of order).
pub(super) index: usize,
pub(super) media: InputMedia,
pub(super) keep_alive: Option<NamedTempFile>,
}
/// Downloads / processes one media item for the upload fallback (see
/// [`send_batch_via_upload`]). Local files are uploaded directly; oversized
/// items fall back to their smaller URL; photos are downscaled/transcoded.
pub(super) async fn prepare_upload_item(
item: MediaItemPayload,
index: usize,
caption: Option<&str>,
) -> Result<PreparedItem, FallbackError> {
// Locally produced files (ugoira / bsky remux MP4): nothing to download
// or shrink — upload the file directly. The send is a multipart upload,
// so the only remaining failure is an upload-cap error, which is
// permanent (a video cannot be re-encoded here).
let media_url = item_url(&item);
if !media_url.starts_with("http://") && !media_url.starts_with("https://") {
let media = media_from_file(
&item,
std::path::PathBuf::from(media_url),
caption,
item.thumbnail_url(),
)
.map_err(|message| FallbackError::Permanent { message })?;
return Ok(PreparedItem {
index,
media,
keep_alive: None,
});
}
// Size check before downloading/uploading: over the cap, use the
// smaller URL instead of the file. Photos are exempt — they are
// downloaded and processed (downscale / PNG→JPEG) before uploading.
let too_large = match x_media::site::media_size(media_url).await {
Ok(Some(size)) => size > MAX_UPLOAD_BYTES,
_ => false,
};
let too_large = too_large && !matches!(item, MediaItemPayload::Photo { .. });
if too_large {
let url = item
.fallback_url()
.ok_or_else(|| FallbackError::Permanent {
message: "media too large".into(),
})?;
let media = media_from_url(&item, url, caption, item.thumbnail_url())
.map_err(|message| FallbackError::Permanent { message })?;
return Ok(PreparedItem {
index,
media,
keep_alive: None,
});
}
match download_to_temp(&item).await {
Ok((file, bytes)) => {
if matches!(item, MediaItemPayload::Photo { .. }) {
// Telegram rejects photos wider+taller than 10000 px combined
// (PHOTO_INVALID_DIMENSIONS): downscale the downloaded file
// before uploading; photos that cannot be brought within the
// limits degrade to the smaller URL. CPU-heavy work runs off
// the async executor thread.
let prep = tokio::task::spawn_blocking(move || photo::prepare_photo(file, &bytes))
.await
.map_err(|e| FallbackError::Permanent {
message: format!("photo worker panicked: {e}"),
})?
.map_err(|message| FallbackError::Permanent { message })?;
match prep {
PhotoPrep::Upload(upload) => {
let path = upload.path().to_path_buf();
let media = media_from_file(&item, path, caption, item.thumbnail_url())
.map_err(|message| FallbackError::Permanent { message })?;
Ok(PreparedItem {
index,
media,
keep_alive: Some(upload),
})
}
PhotoPrep::UseFallback => {
let url = item.fallback_url().ok_or_else(|| FallbackError::Permanent {
message: "photo dimensions exceed Telegram limits and no smaller variant is available"
.into(),
})?;
let media = media_from_url(&item, url, caption, item.thumbnail_url())
.map_err(|message| FallbackError::Permanent { message })?;
Ok(PreparedItem {
index,
media,
keep_alive: None,
})
}
}
} else {
let path = file.path().to_path_buf();
let media = media_from_file(&item, path, caption, item.thumbnail_url())
.map_err(|message| FallbackError::Permanent { message })?;
Ok(PreparedItem {
index,
media,
keep_alive: Some(file),
})
}
}
Err(FallbackError::MediaTooLarge) => {
let url = item
.fallback_url()
.ok_or_else(|| FallbackError::Permanent {
message: "media too large".into(),
})?;
let media = media_from_url(&item, url, caption, item.thumbnail_url())
.map_err(|message| FallbackError::Permanent { message })?;
Ok(PreparedItem {
index,
media,
keep_alive: None,
})
}
Err(e) => Err(e),
}
}
/// Download-and-reupload fallback for one media batch. Files over the upload
/// cap are not downloaded/uploaded; the item falls back to its smaller URL
/// (which Telegram fetches itself). Items are prepared concurrently (bounded)
/// because the downloads are network-bound; the batch is then uploaded in its
/// original order. Returns the fallback-error without the task attached;
/// callers wrap it with the updated task state.
pub(super) async fn send_batch_via_upload(
sender: &dyn MediaSender,
chat_id: i64,
reply_to: i64,
batch: &[MediaItemPayload],
caption: Option<&str>,
task: Task,
) -> Result<Vec<Message>, SendError> {
let sem = std::sync::Arc::new(tokio::sync::Semaphore::new(3));
let mut set = tokio::task::JoinSet::new();
for (i, item) in batch.iter().enumerate() {
let item_caption = if i == 0 {
caption.map(str::to_string)
} else {
None
};
let item = item.clone();
let sem = std::sync::Arc::clone(&sem);
set.spawn(async move {
let _permit = sem.acquire().await.expect("upload semaphore closed");
prepare_upload_item(item, i, item_caption.as_deref()).await
});
}
let mut prepared: Vec<Option<InputMedia>> = (0..batch.len()).map(|_| None).collect();
let mut keep_alive: Vec<NamedTempFile> = Vec::new();
while let Some(joined) = set.join_next().await {
let item = match joined {
Ok(Ok(item)) => item,
// Dropping the JoinSet aborts the remaining prep tasks; their
// temp files are cleaned up on drop (short-circuit like before).
Ok(Err(e)) => return Err(SendError::from_fallback(e, task.clone())),
Err(e) => {
return Err(SendError::Permanent {
message: format!("upload worker panicked: {e}"),
task: Box::new(task),
});
}
};
let PreparedItem {
index,
media,
keep_alive: file_opt,
} = item;
if let Some(file) = file_opt {
keep_alive.push(file);
}
prepared[index] = Some(media);
}
let items: Vec<InputMedia> = prepared
.into_iter()
.map(|m| m.expect("every upload item was prepared"))
.collect();
// `keep_alive` holds the temp files until the group request completes.
let result = sender
.send_media_group(ChatId(chat_id), MessageId(reply_to as i32), items)
.await;
drop(keep_alive);
match result {
Ok(messages) => Ok(messages),
Err(e) => Err(classify_to_send_error(&e, task, "upload failed")),
}
}
+114 -53
View File
@@ -1,12 +1,13 @@
//! Per-chat state with SQLite persistence (table `chat_state` in
//! `data/task_queue.db`, shared with the task queue).
use crate::db::unix_now;
use parking_lot::Mutex;
use rusqlite::params;
use serde::{Deserialize, Serialize};
use std::collections::HashMap;
use std::sync::Arc;
use std::time::{Duration, SystemTime, UNIX_EPOCH};
use std::time::Duration;
#[derive(Serialize, Deserialize, Default, Clone, Debug)]
pub struct ChatData {
@@ -16,8 +17,8 @@ pub struct ChatData {
pub edit_message: HashMap<i64, EditMessage>,
/// name -> HTML template containing "[]"
pub template: HashMap<String, String>,
/// site name (twitter/bsky/pixiv) -> user-supplied caption format with
/// {url} {author} {author_url} {title} {tags} placeholders.
/// site name (twitter/bsky/misskey/pixiv/bilibili) -> user-supplied caption format
/// with {url} {author} {author_url} {title} {content} {tags} placeholders.
pub message_format: HashMap<String, String>,
}
@@ -40,13 +41,6 @@ pub struct ChatStore {
pool: Arc<crate::db::DbPool>,
}
pub fn unix_now() -> i64 {
SystemTime::now()
.duration_since(UNIX_EPOCH)
.map(|d| d.as_secs() as i64)
.unwrap_or(0)
}
impl ChatStore {
/// Wraps the shared DB pool (schema initialized once by
/// [`crate::db::open_store`]; the `chat_state` table lives in the merged
@@ -108,19 +102,22 @@ impl ChatStore {
}
}
/// The per-chat async lock serializing get→mutate→set cycles.
fn lock_for(&self, chat_id: i64) -> Arc<tokio::sync::Mutex<()>> {
self.locks
.lock()
.entry(chat_id)
.or_insert_with(|| Arc::new(tokio::sync::Mutex::new(())))
.clone()
}
/// Serializes a get→mutate→set cycle per chat: concurrent handler tasks
/// (the batch-forward design spawns several per chat) each snapshot the
/// same `ChatData` and last-writer-wins would silently drop mutations,
/// e.g. a second `edit_message` record. The per-chat lock makes the
/// cycle atomic. Returns the closure's result.
pub async fn update<R>(&self, chat_id: i64, f: impl FnOnce(&mut ChatData) -> R) -> R {
let lock = {
let mut locks = self.locks.lock();
locks
.entry(chat_id)
.or_insert_with(|| Arc::new(tokio::sync::Mutex::new(())))
.clone()
};
let lock = self.lock_for(chat_id);
let _guard = lock.lock().await;
let mut data = self.get(chat_id).await;
let r = f(&mut data);
@@ -134,43 +131,45 @@ impl ChatStore {
pub async fn prune_expired(&self, ttl: Duration) -> Vec<(i64, i64)> {
let now = unix_now();
let ttl_secs = ttl.as_secs() as i64;
let mut removed = Vec::new();
// Chats with no live edit records: evicted from the cache (and their
// per-chat lock) so the cache stays bounded to active prompts. The DB
// keeps the row; the next get() reloads it.
let mut evicted_chats = Vec::new();
let changed: Vec<(i64, ChatData)> = {
let mut cache = self.cache.lock();
let mut out = Vec::new();
for (chat_id, data) in cache.iter_mut() {
let keys: Vec<i64> = data.edit_message.keys().copied().collect();
let mut kept = HashMap::new();
for key in keys {
if let Some(entry) = data.edit_message.get(&key) {
if entry.created_at + ttl_secs > now {
kept.insert(key, entry.clone());
} else {
removed.push((*chat_id, key));
}
}
}
if kept.len() != data.edit_message.len() {
// Persist the pruned row (removes expired records from
// the DB too, not just the cache).
data.edit_message = kept;
out.push((*chat_id, data.clone()));
}
if data.edit_message.is_empty() {
evicted_chats.push(*chat_id);
}
}
// Lock order: update() takes the per-chat lock before the cache
// lock, so prune must not hold the cache lock while taking locks.
drop(cache);
out
// Chats that may have an expired record, from a cache snapshot; the
// pruning itself re-reads and writes under the per-chat lock below
// (see the eviction note). Takes no lock of its own, so a chat
// appearing later is simply picked up by the next sweep.
let candidates: Vec<i64> = {
let cache = self.cache.lock();
cache
.iter()
.filter(|(_, data)| {
data.edit_message
.values()
.any(|entry| entry.created_at + ttl_secs <= now)
})
.map(|(chat_id, _)| *chat_id)
.collect()
};
for (chat_id, data) in changed {
self.set(chat_id, &data).await;
let mut removed = Vec::new();
let mut evicted_chats = Vec::new();
for chat_id in candidates {
let lock = self.lock_for(chat_id);
let _guard = lock.lock().await;
let mut data = self.get(chat_id).await;
let before = data.edit_message.len();
data.edit_message.retain(|key, entry| {
if entry.created_at + ttl_secs > now {
return true;
}
removed.push((chat_id, *key));
false
});
if data.edit_message.len() != before {
self.set(chat_id, &data).await;
}
// Chats with no live edit records: evicted from the cache (and
// their per-chat lock) so the cache stays bounded to active
// prompts. The DB keeps the row; the next get() reloads it.
if data.edit_message.is_empty() {
evicted_chats.push(chat_id);
}
}
if !evicted_chats.is_empty() {
let mut cache = self.cache.lock();
@@ -229,4 +228,66 @@ mod tests {
"concurrent get→mutate→set must not drop records"
);
}
fn edit_entry(chat_id: i64, created_at: i64) -> EditMessage {
EditMessage {
url: "https://x.com/u/status/1".into(),
chat_id,
forward_message_ids: vec![9],
template: String::new(),
created_at,
}
}
#[tokio::test]
async fn prune_removes_only_expired_records() {
let dir = tempfile::tempdir().unwrap();
let pool = crate::db::open_store(dir.path().join("p.db").to_str().unwrap()).unwrap();
let store = ChatStore::new(pool);
let now = unix_now();
store
.update(7, |data| {
data.template.insert("t".into(), "[]".into());
data.edit_message.insert(1, edit_entry(7, now - 3600));
data.edit_message.insert(2, edit_entry(7, now));
})
.await;
let removed = store.prune_expired(Duration::from_secs(60)).await;
assert_eq!(removed, vec![(7, 1)]);
let data = store.get(7).await;
assert!(data.edit_message.contains_key(&2), "live record pruned");
assert_eq!(
data.template.get("t").map(String::as_str),
Some("[]"),
"unrelated state lost by the prune"
);
}
#[tokio::test]
async fn prune_eviction_keeps_the_persisted_state() {
// Every record expires → the chat is evicted from the cache; the
// pruned state must already be in the DB when that happens.
let dir = tempfile::tempdir().unwrap();
let pool = crate::db::open_store(dir.path().join("p.db").to_str().unwrap()).unwrap();
let store = ChatStore::new(pool);
store
.update(8, |data| {
data.template.insert("keep".into(), "[]".into());
data.edit_message.insert(1, edit_entry(8, 0));
})
.await;
let removed = store.prune_expired(Duration::from_secs(60)).await;
assert_eq!(removed, vec![(8, 1)]);
let data = store.get(8).await;
assert!(data.edit_message.is_empty());
assert_eq!(
data.template.get("keep").map(String::as_str),
Some("[]"),
"eviction dropped state the DB never received"
);
}
}
+31 -22
View File
@@ -1,8 +1,9 @@
# 架构优化设计:可测试性接缝 + handlers 拆分
> 状态:设计稿(未实施)。目标:把仓库最大的测试空白(`handlers.rs`/`send.rs`
> 发送与分派逻辑)补上可测试接缝,并把 ~1100 行的 handlers 单体拆成模块
> 每个阶段独立提交、独立回滚;全程 fmt / clippy / test 全绿,行为不变。
> 状态:**阶段 A、B、C 已实施**A: `c9e72fd`B: `50206a9` + `ae69d72`C:
> rate_limit 提交);**D 已延迟**——待下次数据库 schema 变化时实施(见 §5)
> 目标:把仓库最大的测试空白(`handlers.rs`/`send.rs` 的发送与分派逻辑)补上
> 可测试接缝,并把 ~1100 行的 handlers 单体拆成模块。
---
@@ -74,26 +75,34 @@ impl MediaSender for Bot { /* 委托现有 teloxide 调用 */ }
**风险**:中。动 `send.rs`/`handlers.rs` 签名(约 15 处调用点),行为不变。
**不做**`main.rs` 的 teloxide 装配不抽象(那是真正的胶水,无测试价值)。
## 4. 阶段 C(可选):主动限流
## 4. 阶段 C:主动限流(已实施)
批量转发时的突发会触发 Telegram 频道限速,现在靠 `RetryAfter → 队列重试` 被动
应对。新增轻量令牌桶(`rate_limit.rs`~50 行):
应对。新增轻量令牌桶(`rate_limit.rs`):
```rust
pub struct TokenBucket { /* capacity, refill_rate, state */ }
pub struct TokenBucket { capacity, refill_per_sec, state: Mutex<State> }
impl TokenBucket {
pub async fn acquire(&self, n: u64) -> Duration; // 等待时长(或 Notify 唤醒)
pub async fn acquire(&self, n: f64); // 按 n 个 token 等待并消费
}
pub fn limiter_for(chat_id: i64) -> Arc<TokenBucket>; // 每频道一个桶
```
- 按频道粒度(`HashMap<ChatId, Arc<TokenBucket>>`),在 `send_media_group`/
`copy_messages` 前置 `acquire`
- 收益:减少 429 → 重试 → 死信;风险低,独立模块。
- 不做的理由(若选不做):当前重试链路已能自愈,容量可按需再加。
- 默认 `CAPACITY = 20``REFILL_PER_SEC = 20/60`(约 20 msg/min);
单次 acquire 可超出容量(记为债务,由后续 refill 偿还)
- 挂点:`MediaSender for Bot``send_media_group`(按 items 数)、
`copy_messages`(按 ids 数)、`send_animation`1 token)前置 `acquire`
MockSender 不受影响(测试不经过限流)。
- 收益:减少 429 → 重试 → 死信;队列重试仍是全局限速的安全网。
- 风险:低,独立模块;`tokio::time`paused-clock 可测)。
## 5. 阶段 D(可选)DB 版本化迁移
## 5. 阶段 DDB 版本化迁移**已延迟**
`schema_init``CREATE TABLE IF NOT EXISTS`,无版本概念。改为:
> ⚠️ **待办提醒**:本阶段**推迟到下次数据库 schema 变化时实施**(给
> `link_cache`/`chat_state`/`tasks` 加列、改结构等)。当前 `schema_init`
> `CREATE TABLE IF NOT EXISTS`,无版本概念;一旦需要迁移已有线上库,必须先落地
> 本方案(`PRAGMA user_version` 迁移链)再改 schema。`db.rs``schema_init`
> 处已留注释指向这里。
```rust
// db.rs
@@ -121,14 +130,14 @@ pub fn migrate(conn: &Connection) -> rusqlite::Result<()> {
- **不引入 DI 框架**:仓库惯例是 LazyLock 静态 + 显式传参,保持。
- **不抽象 main.rs 的 teloxide 装配**
## 7. 提交序列
## 7. 实施记录
| 阶段 | 提交消息(建议) |
|---|---|
| A | `refactor(handlers): split monolithic handlers.rs into modules` |
| B | `refactor(send): introduce MediaSender seam for testable send paths` |
| B+ | `test(send): cover fallback and classification via MockSender` |
| C | `feat(send): add per-chat token bucket rate limiting` |
| D | `refactor(db): versioned schema migrations` |
| 阶段 | 提交 | 说明 |
|---|---|---|
| A | `c9e72fd` | handlers 拆为 `{mod, statics, commands, urls, inline, callback}` |
| B | `50206a9` | `media_sender.rs``trait MediaSender` + `impl for Bot``<Bot as Requester>::` 消歧);send.rs 8 处签名改 `&dyn MediaSender``MockSender` 测试覆盖兜底触发与错误分类(+5 测试) |
| B | `ae69d72` | `AppContext` 注入 `url_media`sender/store/queue/cache),url_media 全链路测试(缓存命中/失效/成功/不支持 URL,+3 测试) |
| C | rate_limit 提交 | `rate_limit.rs` 令牌桶 + 每频道注册表;`MediaSender for Bot` 的 group/copy/animation 前置 `acquire`+3 测试) |
| D | — | **已延迟**:待下次数据库 schema 变化时实施(见 §5) |
每阶段独立合入;A、B 为核心C、D 可选
A、B、C 为核心并已实施;D 在 schema 变更时落地
+4
View File
@@ -10,6 +10,10 @@
## 1. 现状摩擦清单
> ⚠️ 本节记录的是**重构前**的现状:其中的行号、以及 `site/mod.rs` 里的
> `fetch_once` 分派函数(当时的实现)都已不存在,仅作历史记录。当前形态见
> `site/mod.rs``SITES` 注册表——新增站点 = 新模块 + 注册一行。
以现有三站(twitter / bsky / pixiv)为基线,新增第 4 个站点(代号 `example`
今天需要触碰的位置: