fix(fetch): cap one fetch at a 900-second total budget

The per-request CLIENT timeout and the per-chunk DOWNLOAD_IDLE_TIMEOUT both reset on progress, so neither bounded the fetch as a whole: a URL that keeps dripping (a trickle the idle timeout reads as life) could hold one of the eight FETCH_SLOTS for effectively ever, and the queue behind the slots is what then stalls. The attempts loop now runs inside tokio::time::timeout(900 s) — generous for a genuinely large ugoira zip on a honest slow link (minutes), fatal for a drip — and answers Transient with the cache key, so the retry and its jittered backoff take over. The permit is taken outside the timeout: the slot is released on either path. Audit low finding 'url_workers long tail' (process-wide total-deadline option).
This commit is contained in:
2026-09-24 15:38:25 +08:00
parent 7019f34801
commit 5060b760ba
+15
View File
@@ -518,6 +518,13 @@ async fn fetch_with_attempts(url: &str, attempts: u32) -> Result<Option<Fetched>
} }
// Gate every network attempt process-wide (see FETCH_SLOTS). // Gate every network attempt process-wide (see FETCH_SLOTS).
let _permit = FETCH_SLOTS.acquire().await.expect("fetch gate closed"); let _permit = FETCH_SLOTS.acquire().await.expect("fetch gate closed");
let fetch = async {
// One hard ceiling for the whole fetch, backoff naps included (the
// timeout below): the idle timeouts restart on every chunk, so a
// drip-feeding URL could otherwise pin one fetch slot effectively
// forever. Generous for a genuinely large ugoira zip on a slow link
// — minutes, not hours — and the deadline the audit's low finding
// asked for.
for attempt in 0..attempts.max(1) { for attempt in 0..attempts.max(1) {
match site.fetch_from_url(url).await { match site.fetch_from_url(url).await {
Ok(fetched) => { Ok(fetched) => {
@@ -541,6 +548,14 @@ async fn fetch_with_attempts(url: &str, attempts: u32) -> Result<Option<Fetched>
} }
} }
unreachable!("retry loop always returns") unreachable!("retry loop always returns")
};
match tokio::time::timeout(Duration::from_secs(900), fetch).await {
Ok(result) => result,
Err(_) => Err(FetchError::Transient(format!(
"fetch exceeded its 900s total budget [key={}]",
cache_key(url).unwrap_or_else(|| "?".into())
))),
}
} }
/// How long to sleep before retrying `attempt` (0-based) after `err`: the /// How long to sleep before retrying `attempt` (0-based) after `err`: the