Two things the 300 s sweep did the hard way:
- The link cache is pruned by `created_at` (`DELETE FROM link_cache WHERE
created_at < ?`) and had no index on it, so every sweep scanned the whole
table — every post sent inside the TTL window, which is up to a week of
them — while the `url` primary key served none of it. A new migration
(appended; migration 1 is frozen and already shipped) creates the index, and
the upgrade test now asserts it exists after an upgrade.
- `ChatStore::prune_expired` only ever *looked* at chats that had an expired
edit-before-forward record, so a chat with no prompt at all — the common
case: every chat that ever sent a message or ran a command — stayed in the
cache and in the per-chat lock map for the process lifetime. The candidate
set now includes chats holding no records, which is what the eviction below
was written for; the DB keeps the row, so the next use costs one SELECT
(pinned by a new test that also shows the durable settings come back).
Deliberately *not* done: skipping the write in `ChatStore::set` when the state
is unchanged. Comparing against the cached copy would skip a serialize plus a
blocking DB round trip for a no-op update — but every one of the 13 `update`
callers mutates something, so the no-op case is a user repeating an identical
command, and the same comparison would also skip the write that repairs a row
whose earlier write failed. A rare saving against a rare repair, and the write
is what makes the cache a cache rather than a source of truth.
Verified: the new eviction test fails without the candidate change (checked by
reverting it) and passes with it; 118 bot tests and 91 x-media tests pass.
`cargo fmt --check`, `cargo clippy --workspace --all-targets --locked -- -D
warnings` and `cargo test --workspace --locked` clean.
`db.rs`'s two rules — `schema_init` is the version-0 baseline and never gains
a column, `MIGRATIONS` is append-only and never edited — were enforced by
comments only. Both have a silent failure mode across releases, and the worst
one (a baseline edit) makes `open_store` fail with `duplicate column name`,
i.e. a fresh deployment that will not start.
Three tests, no production change:
- A pre-migration database (the historical DDL written out literally, so an
edit to the baseline shows up here instead of being followed) upgrades
through `open_store`: version at the latest, exactly the migrated column
set, rows intact.
- The shipped migration text is frozen and compared entry by entry; the
assertion names the rule when it fires. Appending still passes — that is
the one allowed change.
- A fresh database lands at the latest version, so a deployment that only
ever saw fresh databases is on the same schema as an upgraded one, and
re-opening the same file is a no-op.
Verified by mutation, both confirmed to fail the new tests: editing the
shipped migration (`shipped_migrations_are_frozen`, with the rule in the
message) and adding the column to `schema_init` instead of a migration
(`a_fresh_database_lands_at_the_latest_version`, `duplicate column name:
lease_token`).
Docs synced for this and the previous two items: `main.rs` (startup repair),
`queue.rs` (`runnable_rows`/`replace_payload`), `handlers/` (`urls.rs`'s
repair, `mod.rs`'s `apply_caption_edit`), `db.rs` (the migration tests) and
the untested-modules list (`db.rs` now covered for migrations; the bot-side
live test renamed to match the repo's `--ignored live` filter).
`cargo fmt`, `cargo clippy --workspace --all-targets --locked -- -D
warnings`, `cargo test --workspace --locked` (193 passed, 15 ignored) and
`cargo test -p x-media -- --ignored live` (13) clean.
P2 (hardening) of the retry audit, closing the report's remaining findings.
- Lease fencing. `lease_next` now stamps a random `lease_token`, and every
write-back a worker makes (the 30s heartbeat, `delete_row`, `reschedule`,
`mark_done`) is guarded by it. A lease that expired while its holder was
stalled and was then re-leased used to let *both* holders write the same row:
one duplicated the send, the other silently discarded the new holder's retry
(a 0-row update was not even logged). Now a worker that no longer holds the
lease drops its attempt at the next heartbeat and writes nothing. Reaching
existing databases needed a migration chain, which `db.rs` had been
pre-committed to: `MIGRATIONS` + `migrate` track `PRAGMA user_version`, with
`schema_init` as the version-0 baseline. Verified on a database created
before this change: user_version 0 -> 1, column added, rows intact.
- Dead-letter notifications no longer mislabel an unparsable payload. A row
whose payload no longer deserializes as a `Task` (an older version's shape,
corruption) used to skip the cache invalidation *and* report "Forward failed
permanently" for a send task, because both were derived from the parsed
value. The identity now comes off the raw JSON, so the stale link-cache entry
is dropped and the message names the post.
- Temp files are marked and swept. Every temp file/dir the project creates now
carries `x_media::TEMP_FILE_PREFIX`, and startup removes entries with that
prefix older than an hour — a killed process leaves its downloads (up to
hundreds of MB) behind because no destructor runs, and the age gate keeps the
sweep away from a second instance's in-flight files. Verified live: the log
reports the sweep, an aged leftover goes, a fresh prefixed file and an
unrelated file stay.
- updated_sequence_task: clone the Task and mutate the two fields
instead of rebuilding all 12 by hand (-22 lines; new fields no
longer need a sync here)
- unify unix_now with db::now_f64 (unix_now() = now_f64() as i64),
moved to db.rs next to its clock source
- classify_to_send_error takes the MediaFetchFailure label, folding
the duplicated inline match in send_batch_via_upload (-8 lines)
Every DB operation (queue lease/enqueue, chat_state get/set, link_cache
read/write) used to open a fresh connection — including the busy timeout
and WAL pragma — then close it, on every message, URL job and callback.
Replace with DbPool: a tiny pool (4 connections max, semaphore-bounded
concurrency for backpressure) whose with_conn() method runs the closure on
a pooled connection inside spawn_blocking. Steady-state cost of an
operation is a list pop + semaphore acquire instead of a connection open.
The lease/earliest_run_after queries full-scanned tasks, and the
rollback journal blocked readers behind worker writes. journal_mode=WAL
(persistent, idempotent) plus idx_tasks_pending(status, run_after)
covers both without a schema migration.
Converge the duplicated open_db (open + busy_timeout) and the
spawn_blocking + expect ceremony that every table access repeated
into one db.rs module. ChatStore no longer creates the tasks table
(schema ownership: queue.rs owns tasks, state.rs chat_state,
link_cache.rs link_cache). No schema or behavior change - all
CREATE TABLE statements are byte-identical, IF NOT EXISTS stays
idempotent, so existing data/task_queue.db files need no migration.