fix(retry): stop losing posts to transient failures and broken promises

P0 of the retry audit. The main finding: a Telegram 5xx was classified
Permanent, so one Telegram-side blip dead-lettered the post.

- `classify_request_error`: a server error is retryable again. teloxide sleeps
  10s on a 5xx and then parses the body, so the HTTP status is gone by the
  time the error arrives; it is recognised by shape instead — a JSON
  server-error description, or an `InvalidJson` whose raw body is not JSON
  (a proxy/error page). A JSON body of the wrong shape stays permanent, since
  retrying a type mismatch cannot help. Reproduced end to end: with the old
  classification a fake 502 (HTML body) logged "failed permanently" and
  dead-lettered; now it logs "queued for retry" and the retry delivers.
- The same class of mistake elsewhere: `is_media_fetch_failure` was missing
  `failed to get HTTP url content`, the description single-media URL sends
  answer with, so hotlink-rejected media failed permanently instead of going
  through the reupload fallback.
- `enqueue_retry` now reports whether the row was written, and the callers
  only promise a retry when it was — a failed enqueue (DB write) used to tell
  the user "retrying in Ns" and then deliver nothing, ever.
- A forward that fails retryably now settles the prompt instead of leaving it
  live: the queued row carries the message ids itself, and a live prompt let
  a second Confirm copy the same messages to the channel twice and let Skip
  answer "nothing was forwarded" while the row still delivered.
- A prompt that could not be sent no longer swallows the gated forward
  silently: the chat is told, since nothing would ever forward.
- `scaled_retry_delay` only scales up, so a server-asked `retry_after` above
  the 300s cap is honoured instead of retried early (which earned another 429
  and then dead-lettered the post).
- Download classification: a 4xx media download is permanent (the media is
  gone or refused) while transport errors and 429/5xx retry — previously every
  download error counted as retryable and burned the whole budget. A temp-file
  *write* failure retries too (resource exhaustion clears; a temp dir that
  cannot be created stays permanent).
- Site status mapping: 401/403 are `Blocked` (permanent) rather than
  `Transient`, so a refusal is reported at once instead of after three
  wasted attempts; and a twitter 200 that is not a tweet is no longer
  reported as withheld content (the empty `{}` withheld shape keeps
  `Sensitive`, which is what triggers the auth fallback).
This commit is contained in:
2026-09-20 20:46:18 +08:00
parent 36e5e8afe6
commit 4cf793cd7e
13 changed files with 312 additions and 60 deletions
+39 -8
View File
@@ -277,7 +277,20 @@ pub(crate) async fn post_send_actions(ctx: &AppContext<'_>, task: &Task, message
})
.await;
}
Err(e) => log::error!("failed to send edit prompt: {e}"),
Err(e) => {
log::error!("failed to send edit prompt: {e}");
// Nothing is forwarded until the prompt is confirmed, so a
// prompt that never arrived means this post is never forwarded.
// Tell the chat instead of letting it wait for a prompt that
// will not come.
notify_failure(
ctx.sender,
notify_chat_id,
notify_message_id,
"Could not open the edit-before-forward prompt — nothing was forwarded.",
)
.await;
}
}
return;
}
@@ -301,7 +314,17 @@ pub(crate) async fn post_send_actions(ctx: &AppContext<'_>, task: &Task, message
delay_seconds,
task,
}) => {
enqueue_retry(ctx.task_queue, *task, delay_seconds).await;
// The forward is already committed from the user's side; if it
// cannot be queued, say so rather than going quiet.
if !enqueue_retry(ctx.task_queue, &task, delay_seconds).await {
notify_failure(
ctx.sender,
notify_chat_id,
notify_message_id,
&failure_text(Some(&task), "retry could not be queued"),
)
.await;
}
}
Err(SendError::Permanent { message, .. }) => {
notify_failure(
@@ -316,16 +339,24 @@ pub(crate) async fn post_send_actions(ctx: &AppContext<'_>, task: &Task, message
}
}
/// Enqueues a task for a later attempt (retry / forward resume). When the
/// enqueue itself fails the task can never be sent again, so its keep-alive
/// temp media is released instead of leaking until process exit.
pub(crate) async fn enqueue_retry(queue: &PersistentTaskQueue, task: Task, delay_seconds: f64) {
let payload = serde_json::to_value(&task).expect("task serializes");
/// Enqueues a task for a later attempt (retry / forward resume). Returns
/// whether the retry is actually persisted: when the enqueue itself fails the
/// task can never run again, so its keep-alive temp media is released instead
/// of leaking until process exit — and the caller must not tell the user a
/// retry is coming (nothing would ever deliver it).
pub(crate) async fn enqueue_retry(
queue: &PersistentTaskQueue,
task: &Task,
delay_seconds: f64,
) -> bool {
let payload = serde_json::to_value(task).expect("task serializes");
let run_after = now_f64() + delay_seconds;
if let Err(e) = queue.enqueue(payload, run_after).await {
log::error!("failed to enqueue retry: {e}");
release_keep_alive(&task);
release_keep_alive(task);
return false;
}
true
}
/// Queue entry point: parses the stored task and dispatches.