docs: resync AGENTS.md and the doc comments with the code

AGENTS.md:
- db.rs row claimed a per-store connection pool; there is one shared pool for
  all three tables (statics.rs builds it once).
- retry enqueue moved to send.rs, noted in both handlers rows.
- queue row now names both notifies (workers' + the sweep's).
- Retries bullet documents fetch vs fetch_once.
- test count ~125 -> ~135, the untested-files list no longer claims state.rs
  and handlers.rs are untested, and the live-test inventory mentions the
  token-gated, not-#[ignore]d pixiv download test that makes a local
  `cargo test --workspace` hit the network.
- /bot_dict is admin-only now.

Code docs:
- site/mod.rs: the module doc pointed new sites at `fetch_once` (a name that
  did not exist then and now means a single-attempt fetch) -> `SITES`; the
  cache_key/SITES/Site docs still said "twitter -> bsky -> pixiv" (misskey
  is registered third); RenderData now documents which fields are escaped
  and why url/author_url are not.
- state.rs, callback.rs: drop the pre-misskey site list and the `<name>`
  that rustdoc read as an HTML tag.
- Fixed the remaining rustdoc links/warnings: `cargo doc --workspace
  --no-deps` is now warning-free (was 8).
- docs/site-registry-refactor.md: §1 describes the pre-refactor state; said so.

No behavior change. fmt/clippy clean, 55 + 68 tests pass (the live pixiv
download test flaked on a CDN body timeout, as before).
This commit is contained in:
2026-09-16 23:28:46 +08:00
parent 3fb8421c3a
commit abdc27ed5e
7 changed files with 44 additions and 31 deletions
+22 -14
View File
@@ -2,7 +2,7 @@
//!
//! Dispatch order: twitter → bsky → misskey → pixiv. Each site module
//! exports a `PATTERN`, `enabled()` and `fetch_from_url()`; a future site
//! plugs in by adding one guarded entry in [`fetch_once`].
//! plugs in by adding one guarded entry in `SITES`.
use std::future::Future;
use std::pin::Pin;
@@ -24,9 +24,9 @@ pub use pixiv::PixivError;
/// media list and spoiler flag. Produced by [`fetch`].
#[derive(Debug)]
pub struct Fetched {
/// Canonical URL: x.com/{author}/status/{id} |
/// https://www.pixiv.net/artworks/{id} |
/// https://bsky.app/profile/{handle}/post/{rkey}
/// Canonical URL: `x.com/{author}/status/{id}` |
/// `https://www.pixiv.net/artworks/{id}` |
/// `https://bsky.app/profile/{handle}/post/{rkey}`
pub source_url: String,
/// The exact HTML produced by the site's caption().
pub caption: String,
@@ -46,8 +46,15 @@ pub struct Fetched {
pub(crate) _keep_alive: Option<tempfile::TempDir>,
}
/// Pre-escaped values for `{url} {author} {author_url} {title} {tags}`
/// placeholders in user-supplied caption formats.
/// Values for the `{url} {author} {author_url} {title} {tags}` placeholders in
/// user-supplied caption formats, substituted by [`caption_from_fields`] as
/// HTML text (never as an attribute value).
///
/// `author`, `title` and `tags` come from the site API (post text, display
/// names) and are HTML-escaped at construction. `url` and `author_url` stay
/// raw: they are canonical URLs the adapter builds from numeric ids and
/// API-constrained handles/DIDs, so they carry no escapable character — the
/// bot's `/test` report relies on that when it embeds them.
#[derive(Debug)]
pub(crate) struct RenderData {
pub url: String,
@@ -166,7 +173,7 @@ pub fn caption_from_fields(
/// Stable per-post cache key derived from any supported URL, so variant
/// domains (x.com / twitter.com / fxtwitter.com, mobile, `/photo/N`
/// suffixes) map to the same post. Delegates to each registered site's
/// `cache_key` (dispatch order twitter → bsky → pixiv).
/// `cache_key` (in registry order).
pub fn cache_key(url: &str) -> Option<String> {
SITES.iter().find_map(|site| site.cache_key(url))
}
@@ -275,20 +282,21 @@ pub(crate) fn log_once_ffmpeg_missing() {
}
}
/// Site adapter: one impl per supported site (twitter / bsky / pixiv),
/// registered in [`SITES`]. All site-specific knowledge — URL pattern,
/// Site adapter: one impl per supported site (twitter / bsky / misskey /
/// pixiv), registered in `SITES`. All site-specific knowledge — URL pattern,
/// cache-key format, fetch, retry policy, media-host headers, startup
/// validation — lives in the site module; the central dispatcher only
/// iterates the registry.
///
/// Async methods return a boxed future (see [`SiteFuture`]): `async fn` /
/// Async methods return a boxed future (see `SiteFuture`): `async fn` /
/// RPITIT in traits are not dyn-compatible (verified on rustc 1.95), and
/// `+ Send` is required since URL/queue workers spawn these futures. The
/// site structs are stateless unit structs, so the boxed futures never
/// borrow from `self` beyond the call's scope.
pub trait Site: Send + Sync {
/// Stable site id (`"twitter"` / `"bsky"` / `"pixiv"`): caption-format
/// lookup, cache-key prefixes and the SetFormat whitelist derive from it.
/// Stable site id (`"twitter"` / `"bsky"` / `"misskey"` / `"pixiv"`):
/// caption-format lookup, cache-key prefixes and the SetFormat whitelist
/// derive from it.
fn id(&self) -> &'static str;
/// URL pattern; the dispatcher's first match wins (dispatch order).
fn pattern(&self) -> &'static Regex;
@@ -324,8 +332,8 @@ pub trait Site: Send + Sync {
type SiteFuture<'a, T, E = FetchError> = Pin<Box<dyn Future<Output = Result<T, E>> + Send + 'a>>;
/// The one registry of supported sites, in dispatch order (twitter → bsky →
/// pixiv). Adding a site = new module + one `Box::new(...)` entry here; the
/// bot crate never lists sites itself.
/// misskey → pixiv). Adding a site = new module + one `Box::new(...)` entry
/// here; the bot crate never lists sites itself.
static SITES: LazyLock<Vec<Box<dyn Site>>> = LazyLock::new(|| {
vec![
Box::new(twitter::TwitterSite),
+2 -2
View File
@@ -1,5 +1,5 @@
//! Callback query handling: the edit-before-forward prompt's "forward" and
//! "template|<name>" buttons.
//! Callback query handling: the edit-before-forward prompt's `"forward"` and
//! `"template|<name>"` buttons.
use super::{CHAT_STORE, CONFIG, TASK_QUEUE};
use crate::db::unix_now;
+1 -1
View File
@@ -5,7 +5,7 @@
//! key]. A repeated link is then answered entirely from local state — no
//! re-fetch of the source site, no re-upload — and no media file is stored
//! on disk (the file ids point at Telegram's servers). Entries expire after
//! [`Config::link_cache_ttl`]; a stale entry is dropped lazily on read and
//! `Config::link_cache_ttl`; a stale entry is dropped lazily on read and
//! by the periodic prune in `main`.
use crate::db::now_f64;
+5 -4
View File
@@ -271,10 +271,11 @@ pub async fn invalidate_cache_with(cache: &LinkCache, task: &Task) {
/// Locally produced media files (ugoira MP4, bsky remux MP4) whose temp dirs
/// must stay alive while their task may be retried by the queue. The fetch
/// pipeline hands ownership here via [`x_media::site::Fetched::take_keep_alive`]
/// before the [`Fetched`] is dropped; a queued retry runs after that drop, so
/// without this the local file would be gone by the time the retry sends it.
/// Entries are removed when the task settles (see [`release_keep_alive`]).
/// pipeline hands ownership here via
/// [`x_media::site::Fetched::take_keep_alive`] before that
/// [`x_media::site::Fetched`] is dropped; a queued retry runs after that drop,
/// so without this the local file would be gone by the time the retry sends
/// it. Entries are removed when the task settles (see [`release_keep_alive`]).
pub static KEEP_ALIVE: LazyLock<parking_lot::Mutex<Vec<tempfile::TempDir>>> =
LazyLock::new(|| parking_lot::Mutex::new(Vec::new()));
+2 -2
View File
@@ -17,8 +17,8 @@ pub struct ChatData {
pub edit_message: HashMap<i64, EditMessage>,
/// name -> HTML template containing "[]"
pub template: HashMap<String, String>,
/// site name (twitter/bsky/pixiv) -> user-supplied caption format with
/// {url} {author} {author_url} {title} {tags} placeholders.
/// site name (twitter/bsky/misskey/pixiv) -> user-supplied caption format
/// with {url} {author} {author_url} {title} {tags} placeholders.
pub message_format: HashMap<String, String>,
}