The same fixtures were rebuilt in five test modules: a `CachedPost`
literal in `link_cache.rs`, `handlers/urls.rs` and twice in
`send/mod.rs`, the edit-before-forward prompt in `handlers/mod.rs` and
`handlers/callback.rs`, and a scripted API error in both handler
modules. They now live in `ctx::test_support` next to `TestStores`:
- `cached_photo()` — the canonical cached post (photo + file id at
`https://x.com/u/status/1`, key `twitter:1`); tests mutate the fields
they care about, as the caption-quote test already did.
- `seed_prompt(template, created_at)` + `PROMPT_ID`/`FORWARDED_ID` —
the prompt record, the chat template and the bound forward channel.
The two former copies differed only in which knob the caller set (the
callback tests backdate it for the expiry cases, the reply tests pick
the template), so the union is one helper.
- `api_error(message)` — construction only; each test module keeps its
own message constant, because the wording is what that module's path
answers with (`chat not found` vs `message not found`).
`send/mod.rs`'s `cached_sequence_cache_data()` (which re-extracted the
post out of the task it had just built) is gone: the two settle tests
seed the cache from the same builder the task uses.
No behaviour change: the values are the ones the tests used except
`file_id` (`AgAC-file-id` everywhere, asserted in the link-cache
round-trip) and `sensitive` (the unasserted `true` in the link-cache
fixture), and every test still passes unchanged.
Verified: `cargo fmt`, `cargo clippy --workspace --all-targets --locked
-- -D warnings` and `cargo test --workspace --locked` (180 passed, 14
ignored).
Audit of all 205 tests (five read-only passes plus a line-by-line
re-check). Ten test functions were removed or merged and eight
subsumed assertion blocks trimmed; the suite is down to 180 tests with
no loss of mutation coverage, and four tests that were passing for
nothing now fail when the code they name is broken.
Redundant (deleted or merged):
- twitter: `syndication_text_only_has_no_media` (re-asserts its own
empty fixture), `..._keeps_multibyte_text` (both transforms are
no-ops for that text), `..._strips_trailing_short_link_without_entities`
(same branch as `..._media_short_link`, which now also covers the
real multibyte tweet), `..._regardless_of_index_units` (its
`display_text_range` rationale outlived the function it described).
- pixiv: `test_fetch` (a bare `is_ok()` on the illustration
`download_media_pixiv_original_with_referer` already asserts and
downloads, and the only network touch in a plain `cargo test`),
`startup_validation_only_disables_on_a_definitive_failure` (four rows
that are a subset of the retry-policy table; the `validate()` branch
it was named for is not asserted at all).
- bilibili: `from_item_legacy_draw_shape_still_parses` (its fixture is
the same legacy `draw` shape `from_item_maps_draw_images_and_topic`
builds, with a subset of its assertions).
- site/mod.rs: two `Ok(None)` cases merged into one test.
- urls.rs: `cache_hit_success_keeps_the_cache_entry` (the `/test` test
asserts the same two things under stricter settings), plus a
`assert_ne!` loop that re-states the mapping assertions above it.
- send/mod.rs: `media_group_success_and_forward_ok` (the forward half is
covered by `post_send_forwards_immediately_when_configured`; the
`is_ok()` half cannot see the returned file ids), and two boundary
rows implied by the constant they sit next to.
- commands.rs: the parse tail that `every_documented_invocation_parses`
already covers per README form, and three `debug_report` rows the
escaping test pins with stronger input.
Passing for nothing (now real):
- `truncate_caption_does_not_split_an_html_entity` — the cut lands
inside the entity, so `!contains("&")` never fired; it now asserts
the exact output in both directions and fails when the guard in
`truncate_caption` is deleted (verified).
- `pipeline_resizes_oversized_jpeg` — magic bytes and a non-empty buffer
pass for a copy-through; it now decodes the output's headers and
fails when the JPEG branch skips the resize (verified).
- `live_validate_with_bogus_token_fails` — expected `PixivError::Api`,
which the status check before the body read made unreachable; a bogus
token is a 4xx. Confirmed against the live endpoint: the old
assertion fails with `got Err(Status(400))`, the new one passes.
- bsky `live_fetch_with_photos` — its URL is a text-only post and it had
a byte-identical twin, so no live test pinned media; it now points at a
labelled post with photos and asserts media + the label (live-verified).
Also fixed, found by turning the runtime-sweep test into a real one:
the 30 s lease-expiry sweep recovered crashed rows but never woke a
worker, so a recovered task waited for the next unrelated enqueue (every
worker is parked on `notify` when no row is pending). `recover_update`
now reports its count, `recover_expired` wakes a worker when it changed
something, and `runtime_sweep_recovers_expired_lease` drives the spawned
loop with a paused clock instead of calling the recovery by hand — it
fails on both the missing wake-up and a sweep that recovers nothing.
Verified: `cargo fmt --check`, `cargo clippy --workspace --all-targets
--locked -- -D warnings`, `cargo test --workspace --locked` (180 passed,
14 ignored) and `cargo test -p x-media -- --ignored live` (13 passed).
`cp .env.example .env` is now the documented starting point: the tracked
template carries every variable (grouped required / sites / bot behaviour /
network / webhook / reverse proxy) with the defaults the code would use
anyway, and the compose comment plus both READMEs point at it. The proxy note
is spelled out where it matters — teloxide panics on a blank `TELOXIDE_PROXY`,
and inside a container the proxy host must be `host.docker.internal`.
Two follow-ups the template exposed:
- `RUST_LOG=` (present but blank, which `.env` makes easy) silenced the log
again: "unset" was handled, "empty" was not. A blank value now falls back to
the same default. Verified: blank and unset both produce the full startup
sequence.
- `TWITTER_AUTH_TOKEN` was in the README prose but missing from the env table
(both languages).
Verified: `docker compose --env-file .env.example config -q` resolves, and a
script comparing the compose's `${VAR}` references against the template's keys
finds none missing.
- `docker-compose.yml.example` becomes `docker-compose.yml`, committed as the
real deployment file (the log caps came along) and no longer gitignored.
Every instance value is now `${VAR}` with a `${VAR:-default}` fallback, so
compose reads it from the gitignored `.env` beside the file and the committed
composition needs no per-deployment edit; a variable not listed in a
service's `environment:` never reaches that container. Two knobs the docs
promised but the file lacked are now wired: `DEFAULT_HOST` on nginx-proxy and
`BILIBILI_COOKIE`/`CAPTION_QUOTE_TEXT_CHARS` on the bot, and the healthcheck
follows `WEBHOOK_PORT` instead of a hardcoded 8443. `TELOXIDE_PROXY` is
deliberately *not* passed — a loopback proxy inside a container is the
container itself, and teloxide panics on a blank value — so the compose
comment and both READMEs explain when to add it by hand.
- `BILIBILI_PLAN.md` moves to `docs/BILIBILI_PLAN.md` (nothing referenced it).
- Docs synced: AGENTS.md (the deployment-file row, and `.gitignore` no longer
lists the compose file) and both READMEs (deployment and webhook steps now
say "set it in .env" rather than "edit the compose", plus the proxy trap in
the env table).
P2 (hardening) of the retry audit, closing the report's remaining findings.
- Lease fencing. `lease_next` now stamps a random `lease_token`, and every
write-back a worker makes (the 30s heartbeat, `delete_row`, `reschedule`,
`mark_done`) is guarded by it. A lease that expired while its holder was
stalled and was then re-leased used to let *both* holders write the same row:
one duplicated the send, the other silently discarded the new holder's retry
(a 0-row update was not even logged). Now a worker that no longer holds the
lease drops its attempt at the next heartbeat and writes nothing. Reaching
existing databases needed a migration chain, which `db.rs` had been
pre-committed to: `MIGRATIONS` + `migrate` track `PRAGMA user_version`, with
`schema_init` as the version-0 baseline. Verified on a database created
before this change: user_version 0 -> 1, column added, rows intact.
- Dead-letter notifications no longer mislabel an unparsable payload. A row
whose payload no longer deserializes as a `Task` (an older version's shape,
corruption) used to skip the cache invalidation *and* report "Forward failed
permanently" for a send task, because both were derived from the parsed
value. The identity now comes off the raw JSON, so the stale link-cache entry
is dropped and the message names the post.
- Temp files are marked and swept. Every temp file/dir the project creates now
carries `x_media::TEMP_FILE_PREFIX`, and startup removes entries with that
prefix older than an hour — a killed process leaves its downloads (up to
hundreds of MB) behind because no destructor runs, and the age gate keeps the
sweep away from a second instance's in-flight files. Verified live: the log
reports the sweep, an aged leftover goes, a fresh prefixed file and an
unrelated file stay.
P1 of the retry audit, from the report's "reliability and diagnosis" batch.
- Media downloads no longer share the 30s *total* timeout of metadata
fetches. The size caps allowed 10 MiB (reupload fallback) and 512 MiB
(ugoira frame zip) while the clock allowed 30s, so a slow link made those
posts impossible: `MEDIA_CLIENT` has no total timeout and instead bounds
the response head and every chunk with a 30s *idle* window, which keeps the
stalled-connection protection. Verified against a local probe: the old
policy aborts a 40s download at 30.0s, the new one completes it (2 MiB,
40.1s), and a body that stops delivering still fails after exactly 30s.
- A finished row's write-back is no longer best-effort. `delete_row` failing
left the row `in_progress` with a live lease, so the next sweep flipped it
back to `pending` and re-ran a completed task — a second album, a second
prompt, a second channel copy. Both terminal writes are now retried, and a
delete that still fails falls back to a `done` tombstone that neither the
lease query nor the sweep looks at; reschedule (no safe tombstone: marking
it done would drop the retry silently) logs what the sweep will do.
- bsky and pixiv no longer present a *failed* video conversion as a post with
no media: the remux/ugoira error propagates (pixiv keeps its retry class,
bsky reports Transient), so the user sees the real cause and `fetch` gets
its retries. bsky's "no ffmpeg" case stays a degradation — retrying a
deployment gap cannot help.
- pixiv's token exchange checks the HTTP status before parsing the body, so a
429/5xx from the OAuth endpoint stays retryable instead of becoming a
permanent Api/Json error (via the shared `pixiv_error_is_retryable`), and
startup validation only disables pixiv for a rejected credential — one 503
while the container came up used to turn every later pixiv link into
"pixiv support is disabled".
P0 of the retry audit. The main finding: a Telegram 5xx was classified
Permanent, so one Telegram-side blip dead-lettered the post.
- `classify_request_error`: a server error is retryable again. teloxide sleeps
10s on a 5xx and then parses the body, so the HTTP status is gone by the
time the error arrives; it is recognised by shape instead — a JSON
server-error description, or an `InvalidJson` whose raw body is not JSON
(a proxy/error page). A JSON body of the wrong shape stays permanent, since
retrying a type mismatch cannot help. Reproduced end to end: with the old
classification a fake 502 (HTML body) logged "failed permanently" and
dead-lettered; now it logs "queued for retry" and the retry delivers.
- The same class of mistake elsewhere: `is_media_fetch_failure` was missing
`failed to get HTTP url content`, the description single-media URL sends
answer with, so hotlink-rejected media failed permanently instead of going
through the reupload fallback.
- `enqueue_retry` now reports whether the row was written, and the callers
only promise a retry when it was — a failed enqueue (DB write) used to tell
the user "retrying in Ns" and then deliver nothing, ever.
- A forward that fails retryably now settles the prompt instead of leaving it
live: the queued row carries the message ids itself, and a live prompt let
a second Confirm copy the same messages to the channel twice and let Skip
answer "nothing was forwarded" while the row still delivered.
- A prompt that could not be sent no longer swallows the gated forward
silently: the chat is told, since nothing would ever forward.
- `scaled_retry_delay` only scales up, so a server-asked `retry_after` above
the 300s cap is honoured instead of retried early (which earned another 429
and then dead-lettered the post).
- Download classification: a 4xx media download is permanent (the media is
gone or refused) while transport errors and 429/5xx retry — previously every
download error counted as retryable and burned the whole budget. A temp-file
*write* failure retries too (resource exhaustion clears; a temp dir that
cannot be created stays permanent).
- Site status mapping: 401/403 are `Blocked` (permanent) rather than
`Transient`, so a refusal is reported at once instead of after three
wasted attempts; and a twitter 200 that is not a tweet is no longer
reported as withheld content (the empty `{}` withheld shape keeps
`Sensitive`, which is what triggers the auth fallback).
P2 of the logging plan (the README recipe landed with the code change):
- `docker-compose.yml.example`: one `x-logging` anchor applied to all three
services. json-file grows without limit by default, so a long-running bot
and the proxy in front of it fill the disk; capped at 10m × 3 files.
- The startup `config:` line now reports what the process actually resolved —
the state DB path (a mistyped `DATA_DIR` or a surprising CWD was invisible
until it bit), both TTLs, the caption-quote setting (`off` rather than a bare
`0`) and whether a proxy is configured. The proxy URL is never printed (it
may embed credentials) and admin ids — chat identifiers — stay at `debug`.
Verified: `docker compose config -q` accepts the file, and a scripted fake-API
run shows `caption quote off` / `link cache TTL 3600s` under overrides,
`proxy=yes` with no credential in any line, and the ids at `debug` only.
P0 (foundation) + P1 (diagnostic depth) of the logging plan:
- main.rs initializes the timed builder with a default filter of
`info,hyper_util=warn,reqwest=warn`. Without RUST_LOG nothing was logged at
all (env_logger falls back to `error`), so `docker run --env-file .env` was
silent, and the plain `init` had no timestamps.
- info-and-above lines stop printing user URLs (fetch/send failures, inline
fetch, bsky's remux warnings). The full URL, the message text and the inline
query move to `trace`, so a `debug` log can be handed to someone else.
- Lifecycle lines name the chat and the post: sent/failed/queued plus the
total `ms`, the edit prompt, the channel forward, and every queue line
(`chat=` + `[key=…]` + per-attempt `ms`, dead-letters included).
- Queue work is visible: `x-media`'s fetch line carries its duration (ugoira
encode and HLS remux included), and the 300s sweep reports the pending count
and how overdue the oldest task is — only when the queue is non-empty.
- URL workers are supervised like the queue workers: a panicking worker used
to die silently and shrink the pool for the rest of the process.
- Degradations that still serve the user (cache/state write or read failures,
a failed chat action) are `warn`, not `error`.
Verified against the scripted fake-API harness: unset RUST_LOG logs info with
timestamps, `debug` carries no user URL, `trace` does, a cache-hit send logs
`chat=111 in 5ms`, a failing send queues and dead-letters with chat+key, and
the sweep reports the pending retry.
`/debug` reported `Fetched::caption`, the site's built-in caption, so a
chat's `/set_format` override never showed up in the preview — the command
looked like a no-op, and the `/set_format` success reply now tells users to
preview with `/debug`, which made that advice wrong.
`preview_caption` mirrors the send paths instead: the per-site format
override (`caption_from_fields`, empty → built-in) plus the long-post
quoting. Reported from a live run against the test bot.
teloxide's `split` parser takes exactly one space-separated token per
field, so `/set_format <site> <format>` — a two-token command — never
parsed: `Command::parse` failed, `message_handler` fell through to the URL
flow, and the user got silence. `/clear_cache` without its optional link
failed the same way ("too few arguments"), so clearing everything was
unreachable. Both now use the crate's `parse_arg_remainder` (whole
remainder, trimmed), which is what their executors were already written
against (`split_once(char::is_whitespace)`).
Found by driving the real binary against a scripted fake Bot API: the
`/set_format` replies never appeared while `/set_template` and the other
single-token commands did. `every_documented_invocation_parses` now pins
every documented form, which is what should have caught it.
The placeholder validation and `-` reset added earlier only work now that
the command reaches its executor at all.
`/start` was "Hello!" and `/help` was the bare command list teloxide can
render — no argument syntax, no caption placeholders, no mention that
links only work in private chats. Both now carry that guidance, and the
bot's profile description / short description are set at startup so a
shared link says what the bot does.
`/settings` reports what this chat is configured to do (forward channel,
edit-before-forward, per-site formats, saved templates) to anyone in the
chat — `/bot_dict` is a raw admin-only dump. Templates can be removed
(`/remove_template`, listing the live names on a typo) and the prompt's
keyboard folds 3 per row with a cap: Telegram rejects a keyboard over 100
buttons outright, which would silently drop the whole prompt.
Inline results hand URLs to Telegram, which fetches them without any
site headers — pixiv's pximg.net answers 403 to that, so those items are
skipped instead of shipped broken. `needs_media_headers` answers that
question from the same per-site rule the downloader uses.
Dead-letter and retry notices name the failing post and the cause
(`failure_text`), since "Task failed after retries: task failed after 2
retries" said neither which link it was nor what happened.
Four ways a user could get silence are closed: a registered-but-disabled
site (pixiv without a token) now answers instead of being dropped as an
unsupported link, `/test` on such a link replies instead of doing nothing,
a supported link posted in a group gets a one-line hint (channels stay
silent), and fetch failures name their cause — gone / withheld / source
risk control / site disabled / source down — instead of one generic
sentence. `FetchError::Disabled` carries the "matched but switched off"
answer, which `find_site` used to fold into `Ok(None)`.
A withheld tweet no longer degrades to "no media": without
`TWITTER_AUTH_TOKEN` it stays `Sensitive` so the reply says the media is
age-restricted, and a failed authenticated fallback propagates its own
class instead of masquerading as an empty post (`empty_fetched` is gone).
Long jobs stop looking stalled: `run_with_chat_action` re-sends the chat
action every 4s while the pipeline is pending and the hint switches from
typing to send-photo/video once the media kinds are known. Media groups
go from 9 to Telegram's 10.
`/set_format` rejects unknown `{…}` placeholders (a typo used to be
published verbatim in every caption) and resets with `-`. The
edit-before-forward prompt states its TTL and that Confirm is required,
gains a Skip button, and is rewritten in place to "expired" by the sweep
— an edit, never a new message, so a background timer cannot wake a chat.
A payload written before the title/content split has no `content` field;
`#[serde(default)]` is what keeps it readable, and the cache deletes any
payload it cannot parse — so dropping that default would silently evict
entries rather than degrade them. The test inserts the literal pre-split
JSON and asserts it comes back with its text left in `title` (no
migration: the entry lives one TTL and moving the text would only
reshuffle `/set_format` placeholders until it expires) and its stored
caption untouched.
A post whose text (the split `title` plus `content`, joined by
`site::compose_text`) reaches `CAPTION_QUOTE_TEXT_CHARS` — default 200,
`0` disables — now has that text wrapped in `<blockquote expandable>`
inside its caption, leaving the URL and author line outside the quote.
Applied at the send boundary (`send_media_sequence`, `send_animation` and
the inline answers), where the caption is already truncated and the same
cache snapshot supplies the text, so a fresh send, a link-cache resend
and a queued retry all decide identically. The text is located as what
follows the author link, with the visible prefix accepted as a match
because `truncate_caption` may cut inside it — that keeps the longest
posts, the ones that most need folding, quoted. Captions whose layout
moves the text elsewhere (pixiv's title-inside-a-link, a `/set_format`
that puts `{title}`/`{content}` first) stay unquoted rather than risking
a blockquote nested in a tag, and a caption that already carries one is
never wrapped again.
Telegram measures a caption *after entities parsing*, so the tags cost no
length and the 1024-character limit cannot be breached; retries replay
the unwrapped caption, so a threshold change takes effect immediately.
The edit-before-forward rewrite stays unquoted by design.
Verified against Telegram: a media-group caption built this way comes
back with `caption_entities` `url` @0, `text_link` @50,
`expandable_blockquote` @56 — the quote starts after the author line.
`Fetched.title` carried whatever text the platform had — a tweet's body,
a bilibili dynamic's body, a pixiv artwork's title — which was enough
while x/twitter (no title at all) set the shape. The platforms actually
disagree: pixiv has a title *and* a description, bilibili has an opus
headline *and* a body. Posts now carry both:
- `title`: the platform's title (a pixiv artwork title, a bilibili opus
headline or video card title), empty on text-only platforms;
- `content`: the body (tweet / bsky / misskey text, bilibili dynamic
body, and pixiv's description — fetched for the first time here and
flattened from the app API's HTML to plain text).
`{content}` joins the caption-format placeholders, so a custom
`/set_format` can include a pixiv description. The built-in captions keep
producing byte-identical output: `compose_text` joins the two fields the
same way the single field already was, and bilibili's forward marker
(`//@author:`) now lands in `content` behind the head line's `title`.
`CachedPost.content` is `#[serde(default)]`, so link-cache entries and
queued task payloads written before the split still parse, their text
living in `title`.
An image/text post fetched without `features=itemOpusStyle` comes back in
bilibili's legacy shape, where the post's body and headline are gone
completely — `desc: null`, no `major.opus` — so `title` (and `{title}`)
stayed empty for exactly the posts that do have content
(`opus/1248857553488576532`: legacy `desc: null`, flagged
`major.opus.summary.text = "[doge_金箍]黑白搭配"`). The same flag also
moves the pictures to `major.opus.pics` (key `url`, not `src`).
Text now falls back opus (headline + body) → `desc.text` → archive card
title; media falls back `opus.pics` → `draw.items` → archive cover, so
the legacy shapes keep working if the flag is ever retired.
Verified live: the reported link now yields title "[doge_金箍]黑白搭配"
with its picture; AV dynamics keep their card title; forwards keep
`desc.text` and gain `//@` composition unchanged.
A 视频投稿动态 (`DYNAMIC_TYPE_AV`) carries no body at all: the API
answers `desc: null` and the content is the archive card, so `title`
(and `{title}` in caption formats) stayed empty for the most common
dynamic type. Audited 24 live dynamics: every dynamic that *has* text
(a 图文 post, a forward, a text post) keeps it in
`module_dynamic.desc.text` — only the AV card has none, so the video
title now stands in, mirroring pixiv whose `title` is the artwork title
rather than post text.
Also records two API observations in comments/docs: an id that cannot
exist answers `4101105 请求数据发生错误` (kept on the permanent arm), and
the feed endpoints strip `desc.text` so only the detail endpoint shows
whether a post has text.
Fetch `t.bilibili.com/<id>`, `www.bilibili.com/opus/<id>`,
`t.bilibili.com/h5/dynamic/detail/<id>` and `m.bilibili.com/dynamic/<id>`
through the anonymous `/x/polymer/web-dynamic/v1/detail` endpoint (no
cookie, no WBI signature; the site adds the device cookies
`/x/frontend/finger/spi` hands out, which is what lifts bilibili's
`-352` risk control).
Media: the `major.draw` grid (`.gif` sources become animations, the rest
photos with a downscaled `@518w.jpg` thumbnail used both as preview and
as the oversized fallback), an attached video's cover, and the quoted
dynamic's media for forwards. The video stream itself is not resolved;
`b23.tv` short links stay unmatched (they mostly point at videos, so
matching them would turn a silently ignored link into a failure reply).
`-352`/`-412` map to a retryable error so the queue backs off instead of
dropping the post; a removed dynamic (`500`) is permanent.
Registry-driven, so no bot-side code changes beyond the site lists in the
command replies; found while researching nazurin and
telegram-bili-feed-helper (see BILIBILI_PLAN.md).
Resolve the manifest and lockfile overlap with the rusqlite 0.40 and zip 8.6
bumps that landed on master after this branch was opened:
- crates/xmedia-bot/Cargo.toml: keep rusqlite 0.40 and rand 0.10
- Cargo.lock: re-point x-media/xmedia-bot at rand 0.10.2; teloxide keeps its
own rand 0.8.8 (it requires ^0.8.5), and cargo pruned the now-unused
version_check entry
Verified on the merged tree: cargo metadata --locked, cargo fmt --check,
cargo clippy --workspace --all-targets --locked -- -D warnings,
cargo test --workspace --locked (69 + 73 tests).
rand 0.10 renamed the entry points dependabot's version bump alone cannot
follow: `thread_rng()` -> `rng()`, `gen_range()` -> `random_range()` and the
`Rng` trait -> `RngExt`.
The jitter only needs one float, so use the free function
`rand::random_range(0.2..0.8)` and drop the now-unused trait import instead
of importing `RngExt`.
Verified: cargo fmt --check, cargo clippy --workspace --all-targets -- -D
warnings, cargo test --workspace --locked (retry_delay_seconds_bounds keeps
the 0.2..0.8 jitter window).
Bump both crates (x-media, xmedia-bot) and refresh the lockfile to the
latest semver-compatible releases. No direct dependency or manifest
requirement changed.
- `/test <url>` now runs the ordinary link pipeline and actually sends the
media, but with the chat's post-send actions suppressed: no channel forward,
no edit-before-forward prompt. It is the same code path as a normal link
(same caption/format handling, link cache, retries, dead-letter
notification), so "does this link work?" is answered by the send itself.
- `/debug <url>` keeps what `/test` used to do: fetch and reply with the HTML
parse report, sending/caching/forwarding nothing.
- `urls::url_media` takes a `PostSend` mode (`FromChat` for the URL workers,
`Suppressed` for `/test`); `build_send_task` maps it to the task's
`edit_before_forward`/`forward_channel_id`. Notification ids stay set in both
modes, so a queued retry still reports a dead-letter to the chat.
- `/test` rejects an unsupported URL with the same message the old parse-only
command used (the URL flow would otherwise ignore it silently).
- Report builder renamed `test_parse_report` -> `debug_report` (with the cap
constant), `parse_test_arg` -> `parse_arg_remainder` (now shared by both
commands). README/README.en command tables and AGENTS.md updated; `/help`
descriptions come from the enum.
Tests: +3 (normal flow still honours the chat's settings, `/test` sends with
them suppressed and keeps the cache entry, `build_send_task` mode mapping). The
suppression test was verified to fail when the mode is ignored.
fmt/clippy clean, 73 + 69 tests pass.
- `--locked` on every cargo invocation (ci.yml clippy/test/build, both
Dockerfile builds). The version bump edits Cargo.lock by hand, so a stale
lock must fail loudly instead of being silently re-resolved: CI would
otherwise test a different dependency set than the one committed — and than
the one the released image is built from.
- docker.yml: build the image (no push, no registry login, read-only build
cache) on pull requests touching the build inputs. The Dockerfile's
stub-source machinery, the ffmpeg download and the entrypoint previously
only ran at release time. Also: a release tag must equal both crate versions
before anything is built (the binary carries no version, so `v1.5.1` with
manifests at 1.5.0 used to publish silently wrong tags), `FFMPEG_URL`/
`FFMPEG_SHA256` are taken from repository variables when set, and the
unused `setup-qemu-action` step is gone (single-arch build; the comment says
what arm64 would need).
- ci.yml: `concurrency` cancels superseded runs, `permissions: contents: read`,
`RUST_BACKTRACE=1`, job timeouts, and a release-profile build of the same
package the Dockerfile builds (the profile was otherwise never compiled
before a merge). The `live` job narrows to `-p x-media`: every network- or
secret-gated test lives there, and the bot crate's offline suite already ran
in the `test` job. Timeout is 45 min because the release build is cold on
the first run — a timeout there would kill the job before rust-cache could
save its cache, leaving every later run cold too.
- Actions pinned to commit SHAs (Dependabot keeps them current);
`dtolnay/rust-toolchain` stays on its channel ref by design.
- .github/dependabot.yml: crates (patch bumps grouped), action pins, Docker
base images — the audit gate reports advisories, this is what moves them.
- tokio's `sync` feature is now declared instead of arriving transitively via
teloxide; `.dockerignore` drops docs and markdown.
Verified locally: `cargo fmt --check`, `cargo clippy --workspace
--all-targets --locked`, `cargo test --workspace --locked` (70 + 69 pass),
`cargo build --release --locked` (6m03s cold, the 15.9 MB stripped binary
starts and registers 10 commands), the tag/version gate against both a
matching and a mismatching tag, and YAML parsing of all three workflow files.
send.rs had grown back into the shape handlers.rs was split out of: payload
types, error classification, the download-and-reupload fallback, the senders
and the whole post-send/queue shell in one file. Split by concern, leaving
call sites (`crate::send::x`) unchanged:
- `send/input_media.rs`: payload → `InputFile`/`InputMedia` selection and
`build_media_group` (with its caption-on-first-item rule).
- `send/upload.rs`: the fallback pipeline (download with the upload cap,
photo downscale handoff, smaller-URL fallback, multipart upload).
- `send/post_send.rs`: link-cache write, the `KEEP_ALIVE` registry for locally
produced media, `settle_task`, the post-send actions and the queue entry
points; the parts other modules call are re-exported.
- `send/mod.rs`: payloads, error classification, classification helpers and
the senders themselves, plus the test module.
No behaviour change: 128 + 1326 + 366 + 345 lines, 70 tests still pass.
AGENTS.md updated for the new layout and for `ctx.rs`.
docs/architecture-refactor.md §3 sketched the trait with "按需扩展:
edit_message_caption / delete_message / answer_callback_query …", but only the
five send methods landed, so `callback.rs` and the edit-before-forward caption
swap were stuck on the concrete `Bot` and remained untested (AGENTS.md still
lists callback.rs as untestable).
- `MediaSender` gains `answer_callback_query`, `edit_message_caption` (HTML
parse mode baked in, every caller uses it) and `delete_message`; the mock
records call order plus the texts, captions and answer toasts, so tests can
assert what the user saw.
- `send_message` now returns the sent message id instead of the whole
`Message`: the only consumer of the value is the edit-before-forward prompt
(which keys its record by it), and returning a `Message` forced every mock
to build a teloxide type. `reply`/`reply_html` follow.
- `callback.rs`: the dptree entry only unpacks the update; `handle_callback`
takes plain values + `&AppContext`. `handlers/mod.rs::edit_message_handler`
likewise takes the values the reply carries. Admin/setup APIs
(`get_chat`, `get_chat_administrators`, `get_me`, `set_my_commands`) stay on
the concrete `Bot`: they are not user flows worth a trait.
- The scripted mock moves to `parking_lot::Mutex` (no poisoning unwraps).
Tests: +11 (template button, forward ok/no-channel/retryable, expired+unknown
prompt, caption swap via template, escaping of user text into the caption,
failed swap still consuming the reply, prompt record written by post_send).
fmt/clippy clean, 70 + 69 tests pass.
docs/architecture-refactor.md §3 stopped half-done: `url_media` got an injected
`AppContext`, but `send.rs`'s post-send half kept reaching for the process-wide
`CHAT_STORE`/`TASK_QUEUE`/`LINK_CACHE` statics, so the whole shell after a
successful send (edit-before-forward prompt, channel forward, retry enqueue,
cache write) had no test and no way to get one.
- `ctx.rs` now owns `AppContext` (sender + the three stores + config) with
`from_statics` for production and a `CONTEXT` static for the spawned worker
closures; `handlers/urls.rs` drops its private copy and the duplicated
assembler, and the queue handler/dead-letter callbacks take the context
(main wires them with `CONTEXT`).
- `send_media_sequence`/`send_animation`/`forward_messages`/`post_send_actions`
take `&AppContext`; the cache write goes through the injected cache.
- New `settle_task(ctx, task, Sent|Failed)` is the single place that ends a
task: release its keep-alive temp media, and drop the link-cache entry only
on failure. All five former call sites funnel through it — the earlier
keep-alive leak existed precisely because one of them had to remember.
`invalidate_cache`/`invalidate_cache_with` (static + injected pair, the
latter only existing because of the former) collapse into one private fn.
- `ctx::test_support::TestStores` gives tests a tempdir store set + context;
`handlers/urls.rs` tests use it instead of hand-rolled setup.
Tests: +5 (post-send forward ok / queued / notified, settle Sent/Failed); the
post-send and settle paths were previously untested. fmt/clippy clean,
60 + 69 tests pass.
AGENTS.md:
- db.rs row claimed a per-store connection pool; there is one shared pool for
all three tables (statics.rs builds it once).
- retry enqueue moved to send.rs, noted in both handlers rows.
- queue row now names both notifies (workers' + the sweep's).
- Retries bullet documents fetch vs fetch_once.
- test count ~125 -> ~135, the untested-files list no longer claims state.rs
and handlers.rs are untested, and the live-test inventory mentions the
token-gated, not-#[ignore]d pixiv download test that makes a local
`cargo test --workspace` hit the network.
- /bot_dict is admin-only now.
Code docs:
- site/mod.rs: the module doc pointed new sites at `fetch_once` (a name that
did not exist then and now means a single-attempt fetch) -> `SITES`; the
cache_key/SITES/Site docs still said "twitter -> bsky -> pixiv" (misskey
is registered third); RenderData now documents which fields are escaped
and why url/author_url are not.
- state.rs, callback.rs: drop the pre-misskey site list and the `<name>`
that rustdoc read as an HTML tag.
- Fixed the remaining rustdoc links/warnings: `cargo doc --workspace
--no-deps` is now warning-free (was 8).
- docs/site-registry-refactor.md: §1 describes the pre-refactor state; said so.
No behavior change. fmt/clippy clean, 55 + 68 tests pass (the live pixiv
download test flaked on a CDN body timeout, as before).
- queue: the lease-expiry sweep waited on the workers' `Notify`. `notify_one`
stores a permit, so a sweep wakeup could consume the one meant for a worker,
which then blocked on `notified()` (it only waits when the table looked
empty, i.e. indefinitely) with a due row sitting there. The sweep now has
its own notify, woken only by stop.
- x-media: split fetch's retry loop into `fetch` (3 attempts, unchanged) and
`fetch_once` (1 attempt); inline queries use the latter — the 800ms debounce
plus 1s/2s backoffs were outlasting the answer window of the query.
- send: chunk_media_items now moves items out of the input Vec instead of
requiring `T: Clone` and copying every payload.
fmt/clippy clean, 55 + 69 tests pass.
- commands: /bot_dict dumped the whole chat state to any member of the chat
and could exceed Telegram's 4096-char message limit (the send then failed
and bubbled up as a handler error). It is now admin-only and capped at
MAX_DEBUG_DUMP_CHARS; README, README.en and the /help description updated.
- send: the edit-before-forward template buttons were built from a HashMap
walk, so their order changed between prompts. Now sorted by name.
- link_cache: an unparseable payload (older schema) was reported as a miss
but left in place, re-failing the parse on every later hit; the row is
dropped on read.
- handlers: a link handed to the URL workers after the channel closed
(shutdown) was discarded silently; it is now logged.
- rate_limit: LIMITERS kept one bucket per chat that ever sent media,
forever. The periodic sweep now drops buckets that are idle (refilled to
capacity) and not held by an in-flight sender; acquire's refill was
factored into a shared helper used by the idle check.
Tests: +3 (corrupted row dropped, sorted markup, idle-bucket pruning); the
cache one was verified to fail before the fix. fmt/clippy clean, 55 + 69.
- send: the PHOTO_INVALID_DIMENSIONS marker never matched (the description is
lower-cased, the marker was not), so oversized photos sent by URL were
classified Permanent instead of taking the download-and-downscale fallback.
- inline: the debounce state was one global slot, so a second user's query
cancelled the first user's pending answer entirely; it is now per user.
- state: prune_expired wrote back a stale snapshot without the per-chat lock,
clobbering a concurrent update() (lost edit-message record -> "Expired");
it now re-reads and prunes under the same lock update() uses.
- send/queue: a task dead-lettered on retry exhaustion kept its keep-alive
temp media (ugoira MP4) alive until process exit; dead_letter_notify now
releases it, and enqueue_retry releases when the enqueue itself fails.
Also folds the duplicated retry enqueue in post_send_actions into
send::enqueue_retry (single clock source, single place that releases).
Tests: +6 (marker, per-user debounce x3, keep-alive release, prune contract
x2); the marker and keep-alive cases were verified to fail before the fix.
cargo fmt/clippy clean, 52 + 69 tests pass.
Fourth site adapter: POST /api/notes/show, renote-aware caption and
media normalization, DriveFile type → Illustration/Animated/Video.
Empty thumbnailUrl strings filtered out in thumbnail_for.
- updated_sequence_task: clone the Task and mutate the two fields
instead of rebuilding all 12 by hand (-22 lines; new fields no
longer need a sync here)
- unify unix_now with db::now_f64 (unix_now() = now_f64() as i64),
moved to db.rs next to its clock source
- classify_to_send_error takes the MediaFetchFailure label, folding
the duplicated inline match in send_batch_via_upload (-8 lines)
- photo.rs: chunks_exact(4)/(2) -> as_chunks::<N>().0
(chunks_exact_to_as_chunks, the new lint prefers the
compile-time-checked slice split)
- send.rs: box the Task inside SendError so the error fits the
result_large_err limit (Task is ~400 bytes; the error now moves
through Result as a pointer); unbox with *task at the two
enqueue_retry call sites (handlers/urls.rs, handlers/callback.rs)
cargo clippy --workspace --all-targets is now warning-free; the
remaining proc-macro-error2 future-incompat note is upstream
(teloxide -> aquamarine) and unfixable locally. Full test suite passes.
1.3.0 → 1.4.0: new features (DATA_DIR config, /test blockquote HTML
report) plus the twitter entity-decode, post-send-actions and queue
lease-heartbeat fixes.
Replaces the strip-tags plain-text rendering: the /test reply is now an
HTML message (reply_html helper with ParseMode::Html). Raw fields (url,
source_url, title, author_url, media urls) are escaped, the pre-escaped
render fields are embedded as-is, and the caption is wrapped in
<blockquote>...</blockquote> so the report shows it exactly as it will
render in the sent media caption — escaped text and clickable links
included, no literal &/</> and no raw markup.
The DB file was hardcoded to CWD-relative data/task_queue.db — a footgun
for systemd/cron deployments and a confusing startup failure when the
data/ dir did not exist (SQLite never creates parent dirs).
db_path() now reads DATA_DIR (default data, CWD-relative, unchanged for
local runs and the docker-compose ./data mount) and creates the
directory automatically. README/README.en.md env tables and AGENTS.md
document the new variable.
Runs actions-rust-lang/audit after the offline tests in the test job: a
crate in Cargo.lock with an unfixed security advisory fails the build.
Verified locally against the current lockfile (0 vulnerabilities; the 3
warnings — unmaintained dotenv/proc-macro-error2 and transitive anyhow
unsoundness — do not fail by default).
AGENTS.md claimed user-facing bot strings are Chinese, but every
reply/send_message string in the code is English (Hello!, Send failed,
No media found, Reply to edit message, ...). README stays Chinese;
update both the overview line and the convention line to state the
actual split.
The report's caption line still showed the raw HTML markup
(<a href="...">...</a>). strip_html_tags now drops the tags (keeping
the visible text; the links are already reported via source_url /
author_url) and the remaining entity-encoded text is decoded — the
strip runs on the escaped caption so a tweet text like >^ω^< survives
instead of being eaten as markup. Custom-format captions contain no
tags and pass through unchanged.
The lease was set once to now + LOCK_TTL_SECONDS (120 s) with no
renewal. Tasks that legitimately take longer — slow CDN downloads,
ugoira encodes, rate-limited batch forwards (a 100-message channel copy
waits ~4 min on the per-chat token bucket) — had their lease expire
mid-run; the 30 s expiry sweep flipped the row back to pending and
another worker processed it again, double-sending.
run_with_lease now drives the handler through tokio::select! and
refreshes locked_until every 30 s while it runs. The heartbeat lives in
the same future as the handler, so a panicking worker still lets the
sweep recover the row (no leaked task keeping the lease fresh forever).
A task only reaches the queue after a failed send, so the fresh attempt
never ran post_send_actions (edit-before-forward prompt / channel
forward) — it failed before that point. The old guard skipped
post_send_actions for resumed tasks (batch_index > 0 or sent ids
present), which meant any send that needed a retry after partial
progress silently lost its forward and edit prompt.
post_send_actions is now run unconditionally on a successful queue send;
it executes exactly once, after the whole sequence completed.
AGENTS.md: document the twitter API entity decode and the /test report
HTML-decoded display; add the missing db.rs / media_sender.rs /
rate_limit.rs module rows; fix the statics location (handlers/statics.rs);
refresh test counts (~115, twitter live 5, pixiv api.rs 1, photo heavy
test); versioning convention now includes README.en.md.
README.md / README.en.md: add TELOXIDE_PROXY to the env variable list.
Twitter's syndication and GraphQL APIs return tweet text and display
names pre-escaped for HTML (> < & '); the caption builder
escaped the text again, so sent messages showed literal entities (e.g.
>^ω^< came back as >^ω^<). from_syndication_json now decodes the
API text before storing it — both the syndication path and the
TWITTER_AUTH_TOKEN GraphQL fallback route through it — so the caption
escapes exactly once and renders correctly.
The /test report is a plain-text message but printed the pre-escaped
caption and render fields; it now HTML-decodes them for display so the
report shows the rendered text.
Deleted tweets answer the syndication endpoint with HTTP 200 and a
TweetTombstone (no `errors`, no `id_str`). The body classifier only
knew the `errors` shape, so tombstones fell through to the
`no id_str -> Sensitive` branch and degraded to an empty result,
making the bot reply "No media found" for a deleted tweet.
- extract parse_syndication_body(); tombstone shape -> NotFound
- propagate NotFound from the TWITTER_AUTH_TOKEN GraphQL fallback
instead of swallowing it into empty_fetched
- unit tests for all body classes + live regression test on a real
tombstoned tweet
P0 — level rework + redaction:
- info now carries only lifecycle, per-post business results (sent /
forwarded / copied / template applied), admin actions and anomalies
(upload fallback, retry enqueue; dead-letter stays error).
- Per-request detail moved to debug: message/command logging, URL
extraction, fetching/fetched, link-cache hits, media-group batch sends,
queue processing (enqueue/processing/completed/rescheduled), photo
processing (downscale/transcode), inline queries, sensitive-tweet note.
- Full user-submitted URLs and message text now appear only at debug; at
info and above links are printed via the normalized cache key.
P1 — request correlation:
- handlers::log_key() maps a URL to its normalized post key
(twitter:<id> / pixiv:<id> / bsky:<handle>/<rkey>). The whole lifecycle
of one link (fetch -> send -> cache -> fallback) now logs [key=...], so
multi-worker logs can be correlated by grepping the key.
Convention documented in AGENTS.md.
site::fetch retried every PixivError, so a bad/expired token (403) or a
deleted artwork (404) burned all 3 attempts with backoff against pixiv's
API for nothing. Add PixivError::Status(u16) — the app-API calls now
surface the HTTP status — and retry only the transient classes: network
errors, 429 and 5xx. 4xx / Api (token errors) / Json / NoAuth are
returned immediately. The classification is a pure helper
(fetch_error_is_retryable) with unit tests.
The old stop only set an atomic flag checked between jobs: a worker
blocked in recv() never woke (the channel was never closed), and queued
jobs were neither drained nor abandoned in a defined way despite the
"drains up to 256 jobs" comment. Now stop_url_workers sets the flag,
drops the sender so blocked recv() calls wake with None, and awaits the
worker JoinHandles (each finishes its in-flight job first). main awaits
it inside the existing 30s shutdown timeout.
Smaller/faster production binary (verified: cargo build --release -p
xmedia-bot builds clean with the new profile). panic=abort is
intentionally not set — queue workers and db closures rely on JoinHandle
catching panics, which abort would defeat.
Full English translation of README.md — features, quick start, webhook
deployment (domain + IP-only with acme.sh), env table, command table.
The Chinese README stays the primary one.
A typo in EDIT_MESSAGE_TTL_SECONDS / WEBHOOK_PORT etc. used to silently
fall back to a default, so the bot ran with different behavior than the
operator intended (or failed much later on a bare .expect). Unparseable
values now log a loud warning naming the variable; invalid webhook
settings still surface as a hard .expect in webhook mode.
The stop sequence awaited the queue workers, which can be mid-download
(30s client timeout) or mid-ugoira encode (minutes). A stuck worker would
hold shutdown forever; now the process logs and exits after 30s.
download_to_temp buffered the full bytes, wrote them to a temp file, and
prepare_photo then read the whole file back from disk. The bytes are
already in memory — pass them through (download_to_temp now returns
(file, bytes)) so photo processing never touches the disk for input.
Adds the bytes dependency to xmedia-bot (already in the lock via x-media).
download_media_limited buffered the whole frame zip (cap 512 MB) in
memory before extraction, spiking RAM for large ugoira. New
site::download_media_to_file streams chunks straight to a temp file with
the same Content-Length / stream cap checks, and ugoira_video now opens
the zip from disk inside spawn_blocking. Adds FetchError::Io for local
write failures (hand-rolled error pattern preserved).
The download-and-reupload fallback downloaded each batch item serially,
so a 9-item batch took 9× the slowest download. Items are now prepared
concurrently (bounded to 3 in-flight downloads + photo processing) via a
JoinSet, then the group is uploaded in its original order; the per-item
logic moved into prepare_upload_item. Temp files stay alive until the
group request completes. A failing item still aborts the batch (the
JoinSet drop cancels the remaining prep tasks, as before).
A long tweet text or a pixiv artwork with many tags can exceed Telegram's
1024-char caption cap for HTML parse mode, turning an otherwise fine send
into a permanent 400. truncate_caption() cuts at a char boundary (never
splitting a multi-byte rune or an HTML entity like &) and appends an
ellipsis. Applied in caption_with / caption_from_fields, the link-cache
re-send path and the inline-query captions.
Telegram's sendMediaGroup requires the first item to be a photo when a
group mixes photos and videos; the previous code kept the source-site
order, so a mixed post with a video first (twitter media order is not
guaranteed) would 400 permanently. photos_first() stable-sorts photos
ahead of videos/animations before chunking; within-kind order is kept.
Every DB operation (queue lease/enqueue, chat_state get/set, link_cache
read/write) used to open a fresh connection — including the busy timeout
and WAL pragma — then close it, on every message, URL job and callback.
Replace with DbPool: a tiny pool (4 connections max, semaphore-bounded
concurrency for backpressure) whose with_conn() method runs the closure on
a pooled connection inside spawn_blocking. Steady-state cost of an
operation is a list pop + semaphore acquire instead of a connection open.
The docker workflow only builds/pushes; tests were a local responsibility.
Add .github/workflows/ci.yml with two layers:
- test: cargo fmt --check + cargo clippy --workspace --all-targets -D
warnings + cargo test --workspace (fully offline, no secrets) on every
push/PR, including forks.
- live: the #[ignore]d live-network tests plus the pixiv token-gated
tests, run on schedule / manual dispatch / tag pushes only (fork PRs
cannot read repository secrets), with PIXIV_REFRESH_TOKEN /
TWITTER_AUTH_TOKEN injected and continue-on-error for flaky sites.
Test gating (documented in AGENTS.md):
- live-network tests now carry #[ignore = "live network: ..."] (twitter 3,
bsky 2, pixiv bogus-token 1) and run via -- --ignored live.
- pixiv token tests early-return when PIXIV_REFRESH_TOKEN is absent or
empty (an unset GitHub secret arrives as ""); test_fetch previously
failed without a token.
Also fixes the three clippy assertions_on_constants warnings in send.rs
(required for -D warnings).
A task whose media is a locally produced file (pixiv ugoira MP4, bsky HLS
remux MP4) references a path inside a tempfile TempDir owned by Fetched.
The retry ran after that TempDir was dropped, so the file was already gone
and the retry always failed permanently ("local media file missing") — and
the upload fallback even tried to GET the local path as a URL.
Keep the temp dirs in a process-wide registry (Fetched::take_keep_alive ->
send::KEEP_ALIVE) that is only released when the task settles (sent or
permanently failed); the upload fallback now uploads local files directly
instead of attempting to download them.
Restart-mid-queue still loses the files (documented behavior in
input_file_for) — only the in-process retry path is fixed here.
Telegram fires an inline query on every keystroke and every prefix of a
pasted URL (status/12, status/123, ...) matches the site patterns, so
typing used to trigger a full 3-attempt fetch per keystroke. Answer only
after the query has been stable for 800ms, dedupe repeats through
Telegram's inline cache (explicit cache_time 300), and let a repeat of a
query that produced no answer retry the fetch.
OpenSSL came from two removable defaults: teloxide's `default` feature
(native-tls) and x-media's reqwest default features (default-tls). Switch
both to rustls (webpki-roots) so the binary links no system TLS libs:
- teloxide: default-features=false + rustls + ctrlc_handler (was part of
the removed default)
- reqwest (x-media): default-features=false + rustls-tls
Verified: openssl-sys/native-tls gone from the tree, cargo check clean,
all 5 live twitter/bsky fetches pass over rustls, full container startup
works. The runtime image now needs no libssl.so.3/libcrypto.so.3 or CA
bundle (ffmpeg only processes local files; all downloads go through
reqwest), saving ~8MB.
c40b074 broke the image two ways:
- `cargo clean -p` removes 0 files, so the real sources (host mtimes
older than the step-1 stub build) were never recompiled and the image
shipped the 337KB fn-main stub, exiting 0 on start. Restore the
touch-based rebuild, which forces cargo to see every .rs as newer.
- bookworm-slim does not ship libssl3 despite the old comment; the bot
links OpenSSL via teloxide/reqwest native-tls, so restore the
libssl/libcrypto copies from the builder (same Debian release).