fix: bound the inline state map, test the 300s sweep, add a bot-wide send budget

Three gaps the last audit list named, all in the "resource growth, background
timers and limits nobody watches" class.

**Idle inline-query entries are pruned.** `DebounceStates` had no eviction at
all: one entry per user who ever used inline mode, forever, while the rate
limiter's buckets and the chat store both prune in the 300s sweep. Entries
now carry a `last_seen` stamp and `prune_idle_states()` drops the ones idle
past 300s — the window Telegram caches an inline answer for
(`cache_time(300)`), after which a repeat reaches the bot again and has to be
answered fresh, so the entry would only suppress a fetch the user is waiting
for. The boundary is tested through `prune_idle_at(now, idle_for)` so it does
not depend on ageing a monotonic clock.

**The 300s sweep is a function, and tested.** It was an inline `tokio::spawn`
block: the expiry edit (the only part that talks to Telegram) had no test at
all. It is now `periodic_sweep(sender, chat_store, link_cache, task_queue,
config, stop)`, which also prunes the inline entries, driven in a test with
`start_paused` — the loop's own timer fires the tick, exactly one expired
prompt is rewritten in place, a live one keeps its record and buttons. The
interval is pinned as a constant because no assertion on the edits can see it
(a shorter one produces the same single edit; the paused clock can jump past
the boundary while a tick's DB work is in flight). To make the edit reachable
at all, `edit_message_text` joined the `MediaSender` trait (Bot impl + mock
recording), which is also what keeps `main.rs`'s remaining `Bot` calls
unambiguous. `main.rs` leaves the "untested modules" list except for
startup/shutdown and the dispatcher tree.

**The bot-wide send budget exists.** Telegram throttles a bot in total
(~30 msg/s) as well as per chat; only the per-chat bucket existed, so a batch
forward fanned out over many chats was unguarded and earned 429s the queue
then retried. `acquire_global` charges the same spend against a single shared
bucket at the three paced sites (`send_media_group`, `send_animation`,
`copy_messages`). The unpaced ones (`send_message`, the edits, the toasts) stay
unpaced on purpose: they are one call per action, far below the ceiling, and
pacing a user-visible reply would delay it. Not covered: that the send paths
call it (they need a real `Bot`), which is the same structural gap as the
dispatcher tree.

Also: the startup token-exchange decision is now `startup_validation(result)`
instead of living inside the `Site::validate` future, so "a 5xx while the
container comes up must not disable pixiv" is asserted as a decision — the
message the admin gets plus `enabled()` unchanged. The rejected-credential
half is deliberately not exercised: it calls `disable()`, a process-wide flag
with no reset, and a test touching it would order-couple every other pixiv
test.

Verified: `cargo fmt`, `cargo clippy --workspace --all-targets --locked -- -D
warnings`, `cargo test --workspace --locked` (184 passed, 14 ignored) — plus
mutations, each confirmed to fail the relevant test: the sweep not being
driven on its timer, the interval shortened to 60s, and (earlier) the queue
sweep's missing wake-up. Dropped an empty leftover `crates/x-media/tests/`
directory while there (never tracked by git).
This commit is contained in:
2026-09-21 01:03:15 +08:00
parent 3828d5b483
commit 024dfd50b3
8 changed files with 384 additions and 84 deletions
+40 -7
View File
@@ -1,12 +1,13 @@
//! Per-chat token-bucket rate limiting.
//!
//! Telegram throttles bots that burst past a chat's message budget
//! (roughly 20 messages/min for channels/groups); today the bot absorbs
//! those 429s with queue retries. This limiter smooths the burst *before*
//! it reaches the API: media sends to a chat consume one token per
//! message, refilled at [`REFILL_PER_SEC`], so a batch forward paces itself
//! instead of tripping flood control. The queue retry stays as the safety
//! net for limits this bucket does not model (global per-bot limits etc.).
//! Telegram throttles bots on two budgets: one per chat (roughly 20
//! messages/min for channels/groups) and a bot-wide one (~30 messages per
//! second). Both are smoothed here *before* the burst reaches the API — the
//! per-chat bucket charges one token per message, and [`acquire_global`]
//! charges the same spend against the bot-wide budget, which no per-chat
//! bucket can see (a forward fanned out over many chats spends one token in
//! each and nothing anywhere). The queue retry stays as the safety net for
//! whatever neither bucket models.
use parking_lot::Mutex;
use std::collections::HashMap;
@@ -19,6 +20,12 @@ const CAPACITY: f64 = 20.0;
/// Sustained refill: ~20 messages per minute.
const REFILL_PER_SEC: f64 = 20.0 / 60.0;
/// The bot-wide budget: Telegram allows roughly 30 messages per second for a
/// bot in total, independently of the per-chat limits. Set to the documented
/// ceiling, so it only ever binds on a cross-chat burst.
const GLOBAL_CAPACITY: f64 = 30.0;
const GLOBAL_REFILL_PER_SEC: f64 = 30.0;
struct State {
/// Current token balance; may go negative (debt from an acquire larger
/// than the capacity, repaid by subsequent refills).
@@ -107,6 +114,17 @@ pub fn limiter_for(chat_id: i64) -> Arc<TokenBucket> {
.clone()
}
/// The one bucket every chat shares: Telegram's bot-wide budget.
static GLOBAL_LIMITER: LazyLock<TokenBucket> =
LazyLock::new(|| TokenBucket::new(GLOBAL_CAPACITY, GLOBAL_REFILL_PER_SEC));
/// Waits for `n` messages' worth of the bot-wide budget. Called by the send
/// paths next to their per-chat [`limiter_for`]: at ~30/s it does not bind on
/// a single chat, but a batch fanned out over many chats has no other guard.
pub async fn acquire_global(n: f64) {
GLOBAL_LIMITER.acquire(n).await;
}
/// Drops limiters that are idle (refilled to capacity, so the chat has not
/// sent recently) and are not still held by an in-flight sender. The map
/// would otherwise keep one bucket per chat that ever sent media, forever.
@@ -160,6 +178,21 @@ mod tests {
);
}
#[tokio::test(start_paused = true)]
async fn the_global_budget_is_paced_and_shared() {
// Drain the process-wide budget (no other test touches it: the send
// paths that use it are mocked), then prove the next message waits for
// the refill instead of going out instantly.
acquire_global(GLOBAL_CAPACITY).await;
let start = tokio::time::Instant::now();
acquire_global(1.0).await;
assert!(
start.elapsed() >= Duration::from_secs_f64(1.0 / GLOBAL_REFILL_PER_SEC),
"a fanned-out burst must be paced: elapsed {:?}",
start.elapsed()
);
}
#[tokio::test(start_paused = true)]
async fn prune_idle_drops_full_unheld_buckets_only() {
// Held by this task: kept even at full capacity, a sender has it.