refactor(x-media): split the post title from its content

`Fetched.title` carried whatever text the platform had — a tweet's body,
a bilibili dynamic's body, a pixiv artwork's title — which was enough
while x/twitter (no title at all) set the shape. The platforms actually
disagree: pixiv has a title *and* a description, bilibili has an opus
headline *and* a body. Posts now carry both:

- `title`: the platform's title (a pixiv artwork title, a bilibili opus
  headline or video card title), empty on text-only platforms;
- `content`: the body (tweet / bsky / misskey text, bilibili dynamic
  body, and pixiv's description — fetched for the first time here and
  flattened from the app API's HTML to plain text).

`{content}` joins the caption-format placeholders, so a custom
`/set_format` can include a pixiv description. The built-in captions keep
producing byte-identical output: `compose_text` joins the two fields the
same way the single field already was, and bilibili's forward marker
(`//@author:`) now lands in `content` behind the head line's `title`.

`CachedPost.content` is `#[serde(default)]`, so link-cache entries and
queued task payloads written before the split still parse, their text
living in `title`.
This commit is contained in:
2026-09-18 00:44:49 +08:00
parent 0eb4e5c78d
commit 52184ba6fb
16 changed files with 301 additions and 111 deletions
+1 -1
View File
@@ -24,7 +24,7 @@ Debug command: `/debug <url>` runs the same `x_media::site::fetch` and replies w
The `/test <url>` command runs the ordinary link pipeline (`urls::url_media`) with `PostSend::Suppressed`: the media is sent and cached like any other link, but the chat's `forward_channel_id`/`edit_before_forward` are ignored, so a test never forwards to the channel and never opens the edit prompt (retries and dead-letter notifications behave as usual). Both commands use a custom `parse_arg_remainder` parser (whole remainder, trimmed) because teloxide's built-in `split` parser takes exactly one space-separated token. The `/test <url>` command runs the ordinary link pipeline (`urls::url_media`) with `PostSend::Suppressed`: the media is sent and cached like any other link, but the chat's `forward_channel_id`/`edit_before_forward` are ignored, so a test never forwards to the channel and never opens the edit prompt (retries and dead-letter notifications behave as usual). Both commands use a custom `parse_arg_remainder` parser (whole remainder, trimmed) because teloxide's built-in `split` parser takes exactly one space-separated token.
The `x-media` library: `site::fetch(url)` dispatches through the `SITES` registry (per-site `impl Site`, in order twitter → bsky → misskey → pixiv → bilibili) and returns `Ok(None)` for unmatched URLs. `Fetched { source_url, caption, title, media: Vec<Media>, sensitive, site_id, … }`; `caption_with(format)` substitutes `{url} {author} {author_url} {title} {tags}`. The `x-media` library: `site::fetch(url)` dispatches through the `SITES` registry (per-site `impl Site`, in order twitter → bsky → misskey → pixiv → bilibili) and returns `Ok(None)` for unmatched URLs. `Fetched { source_url, caption, title, content, media: Vec<Media>, sensitive, site_id, … }` (title and content are split per platform: a pixiv artwork's title and description, a bilibili headline and body, and text-only posts whose text is all `content`); `caption_with(format)` substitutes `{url} {author} {author_url} {title} {content} {tags}`.
## Key Directories ## Key Directories
+1 -1
View File
@@ -114,7 +114,7 @@ Telegram only accepts ports 443/80/88/8443.
| `/remove_forward_channel` | Remove the forward channel | | `/remove_forward_channel` | Remove the forward channel |
| `/edit_before_forward` | Toggle "edit before forward": when enabled, the bot posts a prompt after forwarding; replying to it edits the first forwarded message's caption (or taps a template button to apply one) | | `/edit_before_forward` | Toggle "edit before forward": when enabled, the bot posts a prompt after forwarding; replying to it edits the first forwarded message's caption (or taps a template button to apply one) |
| `/set_template <name>` | Reply to a message containing `[]` to save it as a named template; `[]` is replaced by the original post link when forwarding (used with "edit before forward") | | `/set_template <name>` | Reply to a message containing `[]` to save it as a named template; `[]` is replaced by the original post link when forwarding (used with "edit before forward") |
| `/set_format <site> <format>` | Customize the caption format for one site. Sites: `twitter` / `bsky` / `pixiv` / `misskey` / `bilibili`. Placeholders: `{url}` `{author}` `{author_url}` `{title}` `{tags}` | | `/set_format <site> <format>` | Customize the caption format for one site. Sites: `twitter` / `bsky` / `pixiv` / `misskey` / `bilibili`. Placeholders: `{url}` `{author}` `{author_url}` `{title}` `{content}` `{tags}` |
| `/clear_cache [link]` | Clear the link cache (admin only); with a link only that entry, otherwise everything | | `/clear_cache [link]` | Clear the link cache (admin only); with a link only that entry, otherwise everything |
| `/bot_dict` | Show the current chat state (debugging; admin only) | | `/bot_dict` | Show the current chat state (debugging; admin only) |
| `/test <link>` | Parse a link and send its media; no channel forward, no edit-before-forward prompt (send only) | | `/test <link>` | Parse a link and send its media; no channel forward, no edit-before-forward prompt (send only) |
+1 -1
View File
@@ -114,7 +114,7 @@ Telegram 只接受 443/80/88/8443 端口。
| `/remove_forward_channel` | 取消转发频道 | | `/remove_forward_channel` | 取消转发频道 |
| `/edit_before_forward` | 开关「转发前编辑」:开启后,转发成功后 bot 会发一条提示消息,回复它可修改第一条转发消息的 caption(或点击模板按钮套用模板) | | `/edit_before_forward` | 开关「转发前编辑」:开启后,转发成功后 bot 会发一条提示消息,回复它可修改第一条转发消息的 caption(或点击模板按钮套用模板) |
| `/set_template <名称>` | 回复一条含 `[]` 的消息,将其保存为命名模板;转发时 `[]` 会被替换为原帖链接(配合「转发前编辑」使用) | | `/set_template <名称>` | 回复一条含 `[]` 的消息,将其保存为命名模板;转发时 `[]` 会被替换为原帖链接(配合「转发前编辑」使用) |
| `/set_format <站点> <格式>` | 自定义某站点的 caption 格式。站点:`twitter` / `bsky` / `pixiv` / `misskey` / `bilibili`。占位符:`{url}` `{author}` `{author_url}` `{title}` `{tags}` | | `/set_format <站点> <格式>` | 自定义某站点的 caption 格式。站点:`twitter` / `bsky` / `pixiv` / `misskey` / `bilibili`。占位符:`{url}` `{author}` `{author_url}` `{title}` `{content}` `{tags}` |
| `/clear_cache [链接]` | 清空链接缓存(仅管理员);带链接只清该条,否则清空全部 | | `/clear_cache [链接]` | 清空链接缓存(仅管理员);带链接只清该条,否则清空全部 |
| `/bot_dict` | 查看当前聊天状态(调试用;仅管理员) | | `/bot_dict` | 查看当前聊天状态(调试用;仅管理员) |
| `/test <链接>` | 解析链接并发送媒体;不转发到频道、不弹转发前编辑提示(仅发送) | | `/test <链接>` | 解析链接并发送媒体;不转发到频道、不弹转发前编辑提示(仅发送) |
+66 -53
View File
@@ -29,7 +29,7 @@
use super::model; use super::model;
use crate::media::Media; use crate::media::Media;
use crate::site::{FetchError, Fetched, RenderData, Site, SiteFuture}; use crate::site::{FetchError, Fetched, RenderData, Site, SiteFuture, compose_text};
use html_escape::{encode_double_quoted_attribute, encode_text}; use html_escape::{encode_double_quoted_attribute, encode_text};
use regex::Regex; use regex::Regex;
use std::sync::LazyLock; use std::sync::LazyLock;
@@ -270,7 +270,8 @@ impl From<model::Item> for Fetched {
let url = format!("https://www.bilibili.com/opus/{}", item.id_str); let url = format!("https://www.bilibili.com/opus/{}", item.id_str);
let author = author_name(&item).to_string(); let author = author_name(&item).to_string();
let author_url = author_url(&item).unwrap_or_else(|| url.clone()); let author_url = author_url(&item).unwrap_or_else(|| url.clone());
let text = text_of(&item); let (title, content) = text_parts(&item);
let text = compose_text(&title, &content);
let tags = topic_name(&item).to_string(); let tags = topic_name(&item).to_string();
let caption = caption(&url, &author_url, &author, &text); let caption = caption(&url, &author_url, &author, &text);
@@ -279,7 +280,8 @@ impl From<model::Item> for Fetched {
Fetched { Fetched {
source_url: url.clone(), source_url: url.clone(),
caption, caption,
title: text.clone(), title: title.clone(),
content: content.clone(),
media, media,
sensitive: false, sensitive: false,
site_id: "bilibili", site_id: "bilibili",
@@ -287,7 +289,8 @@ impl From<model::Item> for Fetched {
url, url,
author: encode_text(&author).into_owned(), author: encode_text(&author).into_owned(),
author_url, author_url,
title: encode_text(&text).into_owned(), title: encode_text(&title).into_owned(),
content: encode_text(&content).into_owned(),
tags: encode_text(&tags).into_owned(), tags: encode_text(&tags).into_owned(),
}), }),
_keep_alive: None, _keep_alive: None,
@@ -338,37 +341,33 @@ fn archive_title(item: &model::Item) -> Option<&str> {
.filter(|title| !title.trim().is_empty()) .filter(|title| !title.trim().is_empty())
} }
/// The dynamic's own words, richest source first: the opus document /// The dynamic's own words, richest source first: the opus document's
/// (headline plus body) → `module_dynamic.desc.text` → the attached video's /// headline and body → the legacy body (`module_dynamic.desc.text`) → the
/// card title. /// attached video's card title, which is the text of a 视频投稿动态 because
/// such a post has no body anywhere.
/// ///
/// The opus shape is what makes ordinary 图文 posts readable at all — their /// The opus shape is what makes ordinary 图文 posts readable at all — their
/// legacy serialization has no text — while a 视频投稿动态 has no body /// legacy serialization has no text.
/// anywhere and is represented by its card title (mirroring pixiv, whose fn own_parts(item: &model::Item) -> (String, String) {
/// `title` is the artwork title rather than post text).
fn own_text(item: &model::Item) -> String {
if let Some(opus) = opus(item) { if let Some(opus) = opus(item) {
let title = opus.title.as_deref().unwrap_or_default().trim(); let title = opus.title.as_deref().unwrap_or_default().trim().to_string();
let body = opus let body = opus
.summary .summary
.as_ref() .as_ref()
.map(|summary| summary.text.trim()) .map(|summary| summary.text.trim().to_string())
.unwrap_or_default(); .unwrap_or_default();
let text = match (title.is_empty(), body.is_empty()) { if !title.is_empty() || !body.is_empty() {
(false, false) => format!("{title}\n{body}"), return (title, body);
(false, true) => title.to_string(),
(true, false) => body.to_string(),
(true, true) => String::new(),
};
if !text.is_empty() {
return text;
} }
} }
let body = desc_text(item).trim(); let body = desc_text(item).trim();
if !body.is_empty() { if !body.is_empty() {
return body.to_string(); return (String::new(), body.to_string());
} }
archive_title(item).unwrap_or_default().to_string() (
archive_title(item).unwrap_or_default().to_string(),
String::new(),
)
} }
fn topic_name(item: &model::Item) -> &str { fn topic_name(item: &model::Item) -> &str {
@@ -380,28 +379,29 @@ fn topic_name(item: &model::Item) -> &str {
.unwrap_or_default() .unwrap_or_default()
} }
/// The post's text: its own plus the quoted original's when it is a forward, /// The post's title and body: the dynamic's own, with the quoted original's
/// marked the way bilibili's web UI does (`//@author:`). /// text appended to the body the way bilibili's web UI shows forwards
fn text_of(item: &model::Item) -> String { /// (`//@author:`).
let own = own_text(item); fn text_parts(item: &model::Item) -> (String, String) {
let (title, mut content) = own_parts(item);
let Some(orig) = item.orig.as_deref() else { let Some(orig) = item.orig.as_deref() else {
return own; return (title, content);
}; };
let orig_text = own_text(orig); let (orig_title, orig_content) = own_parts(orig);
let orig_text = compose_text(&orig_title, &orig_content);
if orig_text.is_empty() { if orig_text.is_empty() {
return own; return (title, content);
} }
let name = author_name(orig); let name = author_name(orig);
let mut text = own; if !content.is_empty() {
if !text.is_empty() { content.push('\n');
text.push('\n');
} }
if name.is_empty() { if name.is_empty() {
text.push_str(&orig_text); content.push_str(&orig_text);
} else { } else {
text.push_str(&format!("//@{name}:\n{orig_text}")); content.push_str(&format!("//@{name}:\n{orig_text}"));
} }
text (title, content)
} }
/// The dynamic's media: its own grid (or video cover), falling back to the /// The dynamic's media: its own grid (or video cover), falling back to the
@@ -592,7 +592,8 @@ mod tests {
"https://www.bilibili.com/opus/1245284537985925159" "https://www.bilibili.com/opus/1245284537985925159"
); );
assert_eq!(fetched.site_id, "bilibili"); assert_eq!(fetched.site_id, "bilibili");
assert_eq!(fetched.title, "新歌上线"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "新歌上线");
assert!(!fetched.sensitive); assert!(!fetched.sensitive);
assert_eq!(fetched.media.len(), 1); assert_eq!(fetched.media.len(), 1);
match &fetched.media[0] { match &fetched.media[0] {
@@ -627,6 +628,7 @@ mod tests {
Some(( Some((
"索尼音乐中国", "索尼音乐中国",
"https://space.bilibili.com/486906719", "https://space.bilibili.com/486906719",
"",
"新歌上线", "新歌上线",
"音乐" "音乐"
)) ))
@@ -696,9 +698,10 @@ mod tests {
}; };
let fetched = parse(item); let fetched = parse(item);
assert_eq!(fetched.title, "每个人的青春里,都有一首 A-Lin");
assert_eq!( assert_eq!(
fetched.title, fetched.content,
"每个人的青春里,都有一首 A-Lin\n那些曾经陪你失恋的歌\n\n【活动】详情见正文" "那些曾经陪你失恋的歌\n\n【活动】详情见正文"
); );
assert!( assert!(
fetched.caption.contains("那些曾经陪你失恋的歌"), fetched.caption.contains("那些曾经陪你失恋的歌"),
@@ -726,8 +729,8 @@ mod tests {
// A body without a headline, and a headline without a body, both // A body without a headline, and a headline without a body, both
// stand alone rather than rendering an empty line. // stand alone rather than rendering an empty line.
for (title, body, expected) in [ for (title, body, expected) in [
(None, "只有正文", "只有正文"), (None, "只有正文", ("", "只有正文")),
(Some("只有标题"), "", "只有标题"), (Some("只有标题"), "", ("只有标题", "")),
] { ] {
let major = serde_json::json!({ let major = serde_json::json!({
"type": "MAJOR_TYPE_OPUS", "type": "MAJOR_TYPE_OPUS",
@@ -735,7 +738,8 @@ mod tests {
}); });
let mut json = item_json(major, ""); let mut json = item_json(major, "");
json["modules"]["module_dynamic"]["desc"] = serde_json::Value::Null; json["modules"]["module_dynamic"]["desc"] = serde_json::Value::Null;
assert_eq!(parse(json).title, expected); let fetched = parse(json);
assert_eq!((fetched.title.as_str(), fetched.content.as_str()), expected);
} }
} }
@@ -748,7 +752,8 @@ mod tests {
draw_item("http://i0.hdslb.com/bfs/new_dyn/l.jpg"), draw_item("http://i0.hdslb.com/bfs/new_dyn/l.jpg"),
"legacy 正文", "legacy 正文",
)); ));
assert_eq!(fetched.title, "legacy 正文"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "legacy 正文");
assert_eq!(fetched.media.len(), 1); assert_eq!(fetched.media.len(), 1);
assert_eq!( assert_eq!(
fetched.media[0].url(), fetched.media[0].url(),
@@ -811,6 +816,7 @@ mod tests {
fetched.render_fields().unwrap().2, fetched.render_fields().unwrap().2,
"GTX760游戏性能测试,二手显卡尚能战否?" "GTX760游戏性能测试,二手显卡尚能战否?"
); );
assert_eq!(fetched.content, "");
assert_eq!(fetched.media.len(), 1); assert_eq!(fetched.media.len(), 1);
// Forwarding a video dynamic: the quoted card title lands after the // Forwarding a video dynamic: the quoted card title lands after the
@@ -819,8 +825,9 @@ mod tests {
forward["modules"]["module_dynamic"]["desc"] = serde_json::Value::Null; forward["modules"]["module_dynamic"]["desc"] = serde_json::Value::Null;
forward["orig"] = item; forward["orig"] = item;
let fetched = parse(forward); let fetched = parse(forward);
assert_eq!(fetched.title, "");
assert_eq!( assert_eq!(
fetched.title, fetched.content,
"//@索尼音乐中国:\nGTX760游戏性能测试,二手显卡尚能战否?" "//@索尼音乐中国:\nGTX760游戏性能测试,二手显卡尚能战否?"
); );
assert_eq!( assert_eq!(
@@ -851,7 +858,8 @@ mod tests {
fetched.media[0].url(), fetched.media[0].url(),
"https://i0.hdslb.com/bfs/new_dyn/o.jpg" "https://i0.hdslb.com/bfs/new_dyn/o.jpg"
); );
assert_eq!(fetched.title, "转发理由\n//@A-SOUL_Official:\n原动态正文"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "转发理由\n//@A-SOUL_Official:\n原动态正文");
// The forwarder stays the author; the quote appears in the text. // The forwarder stays the author; the quote appears in the text.
assert!( assert!(
fetched.caption.contains("索尼音乐中国"), fetched.caption.contains("索尼音乐中国"),
@@ -870,7 +878,8 @@ mod tests {
fn from_item_without_major_has_no_media() { fn from_item_without_major_has_no_media() {
let fetched = parse(item_json(serde_json::Value::Null, "只有文字")); let fetched = parse(item_json(serde_json::Value::Null, "只有文字"));
assert!(fetched.media.is_empty()); assert!(fetched.media.is_empty());
assert_eq!(fetched.title, "只有文字"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "只有文字");
} }
#[test] #[test]
@@ -879,7 +888,8 @@ mod tests {
draw_item("http://i0.hdslb.com/bfs/new_dyn/a.jpg"), draw_item("http://i0.hdslb.com/bfs/new_dyn/a.jpg"),
"<b>\"x\" & y</b>", "<b>\"x\" & y</b>",
)); ));
assert_eq!(fetched.title, "<b>\"x\" & y</b>"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "<b>\"x\" & y</b>");
// `encode_text` escapes markup only; a bare quote is text, not an // `encode_text` escapes markup only; a bare quote is text, not an
// attribute delimiter, and stays as-is. // attribute delimiter, and stays as-is.
assert!( assert!(
@@ -887,8 +897,8 @@ mod tests {
"{}", "{}",
fetched.caption fetched.caption
); );
let (_, _, title, _) = fetched.render_fields().unwrap(); let (_, _, _, content, _) = fetched.render_fields().unwrap();
assert_eq!(title, "&lt;b&gt;\"x\" &amp; y&lt;/b&gt;"); assert_eq!(content, "&lt;b&gt;\"x\" &amp; y&lt;/b&gt;");
} }
#[test] #[test]
@@ -980,20 +990,22 @@ mod tests {
return; return;
}; };
assert!(fetched.media.is_empty()); assert!(fetched.media.is_empty());
assert!(!fetched.title.trim().is_empty()); assert!(fetched.title.is_empty());
assert!(!fetched.content.trim().is_empty());
} }
/// Regression for the reported case: this opus post's legacy /// Regression for the reported case: this opus post's legacy
/// serialization has `desc: null`, so without the `itemOpusStyle` /// serialization has `desc: null`, so without the `itemOpusStyle`
/// request it parsed with an empty title. /// request it parsed with no text at all.
#[tokio::test] #[tokio::test]
#[ignore = "live network: requires outbound HTTPS to api.bilibili.com"] #[ignore = "live network: requires outbound HTTPS to api.bilibili.com"]
async fn live_fetch_opus_dynamic_has_title_and_text() { async fn live_fetch_opus_dynamic_has_content() {
let Some(fetched) = live_fetch("https://www.bilibili.com/opus/1248857553488576532").await let Some(fetched) = live_fetch("https://www.bilibili.com/opus/1248857553488576532").await
else { else {
return; return;
}; };
assert_eq!(fetched.title, "[doge_金箍]黑白搭配"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "[doge_金箍]黑白搭配");
assert!(fetched.caption.ends_with("黑白搭配"), "{}", fetched.caption); assert!(fetched.caption.ends_with("黑白搭配"), "{}", fetched.caption);
let urls: Vec<&str> = fetched.media.iter().map(|m| m.url()).collect(); let urls: Vec<&str> = fetched.media.iter().map(|m| m.url()).collect();
assert_eq!(urls.len(), 1, "{urls:?}"); assert_eq!(urls.len(), 1, "{urls:?}");
@@ -1008,6 +1020,7 @@ mod tests {
return; return;
}; };
assert!(!fetched.title.trim().is_empty(), "{fetched:?}"); assert!(!fetched.title.trim().is_empty(), "{fetched:?}");
assert!(fetched.content.is_empty(), "{fetched:?}");
assert!( assert!(
fetched.caption.contains(&fetched.title), fetched.caption.contains(&fetched.title),
"{}", "{}",
+7 -3
View File
@@ -330,13 +330,16 @@ impl From<Post> for Fetched {
url: url.clone(), url: url.clone(),
author: encode_text(&post.author).into_owned(), author: encode_text(&post.author).into_owned(),
author_url: author_url.clone(), author_url: author_url.clone(),
title: encode_text(&post.text).into_owned(), // A post has no title: its text is all content.
title: String::new(),
content: encode_text(&post.text).into_owned(),
tags: String::new(), tags: String::new(),
}); });
Fetched { Fetched {
source_url: url, source_url: url,
caption: post.caption(), caption: post.caption(),
title: post.text.clone(), title: String::new(),
content: post.text.clone(),
media: post.media, media: post.media,
sensitive: post.sensitive, sensitive: post.sensitive,
site_id: "bsky", site_id: "bsky",
@@ -410,7 +413,8 @@ mod tests {
fetched.source_url, fetched.source_url,
"https://bsky.app/profile/user.bsky.social/post/3xxxx" "https://bsky.app/profile/user.bsky.social/post/3xxxx"
); );
assert_eq!(fetched.title, "hello <world>"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "hello <world>");
assert_eq!(fetched.media.len(), 1); assert_eq!(fetched.media.len(), 1);
assert!(!fetched.sensitive); assert!(!fetched.sensitive);
// display_name absent -> empty fallback // display_name absent -> empty fallback
+17 -11
View File
@@ -123,21 +123,23 @@ impl From<model::Note> for Fetched {
let cw = content.cw.as_deref().unwrap_or_default(); let cw = content.cw.as_deref().unwrap_or_default();
// Notes carry hashtags inline in the text (no structured tags array); // Notes carry hashtags inline in the text (no structured tags array);
// a CW note gets the marker prefixed so recipients see the spoiler. // a CW note gets the marker prefixed so recipients see the spoiler.
let mut title = cw.to_string(); let mut text = cw.to_string();
if !cw.is_empty() && !title.ends_with(' ') { if !cw.is_empty() && !text.ends_with(' ') {
title.push(' '); text.push(' ');
} }
title.push_str(content.text.as_deref().unwrap_or_default().trim()); text.push_str(content.text.as_deref().unwrap_or_default().trim());
let title = title.trim().to_string(); let text = text.trim().to_string();
let caption = caption(&url, &author_url, &author, &title); let caption = caption(&url, &author_url, &author, &text);
let sensitive = content.cw.is_some() || content.files.iter().any(|f| f.is_sensitive); let sensitive = content.cw.is_some() || content.files.iter().any(|f| f.is_sensitive);
let media: Vec<Media> = content.files.iter().filter_map(media_from_file).collect(); let media: Vec<Media> = content.files.iter().filter_map(media_from_file).collect();
Fetched { Fetched {
source_url: url.clone(), source_url: url.clone(),
caption, caption,
title: title.clone(), // A note has no title: its text (CW marker included) is content.
title: String::new(),
content: text.clone(),
media, media,
sensitive, sensitive,
site_id: "misskey", site_id: "misskey",
@@ -145,7 +147,8 @@ impl From<model::Note> for Fetched {
url, url,
author: encode_text(&author).into_owned(), author: encode_text(&author).into_owned(),
author_url: author_url.clone(), author_url: author_url.clone(),
title: encode_text(&title).into_owned(), title: String::new(),
content: encode_text(&text).into_owned(),
tags: String::new(), tags: String::new(),
}), }),
_keep_alive: None, _keep_alive: None,
@@ -257,7 +260,8 @@ mod tests {
"https://misskey.io/notes/aotihl10lqrs015s" "https://misskey.io/notes/aotihl10lqrs015s"
); );
assert_eq!(fetched.site_id, "misskey"); assert_eq!(fetched.site_id, "misskey");
assert_eq!(fetched.title, "hello"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "hello");
assert!(fetched.sensitive); assert!(fetched.sensitive);
assert_eq!(fetched.media.len(), 1); assert_eq!(fetched.media.len(), 1);
match &fetched.media[0] { match &fetched.media[0] {
@@ -308,7 +312,8 @@ mod tests {
note["text"] = serde_json::json!("body"); note["text"] = serde_json::json!("body");
let fetched: Fetched = note_json(note).into(); let fetched: Fetched = note_json(note).into();
assert!(fetched.sensitive); assert!(fetched.sensitive);
assert_eq!(fetched.title, "spoiler body"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "spoiler body");
} }
#[test] #[test]
@@ -341,7 +346,8 @@ mod tests {
} }
}); });
let fetched: Fetched = note_json(note).into(); let fetched: Fetched = note_json(note).into();
assert_eq!(fetched.title, "inner text"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "inner text");
assert_eq!(fetched.media.len(), 1); assert_eq!(fetched.media.len(), 1);
// The source URL still points at the renote shell the user posted. // The source URL still points at the renote shell the user posted.
assert_eq!( assert_eq!(
+58 -18
View File
@@ -21,8 +21,8 @@ pub mod twitter;
pub use pixiv::PixivError; pub use pixiv::PixivError;
/// The result of fetching a post: canonical URL, HTML caption, raw text, /// The result of fetching a post: canonical URL, HTML caption, the post's
/// media list and spoiler flag. Produced by [`fetch`]. /// title and body, media list and spoiler flag. Produced by [`fetch`].
#[derive(Debug)] #[derive(Debug)]
pub struct Fetched { pub struct Fetched {
/// Canonical URL: `x.com/{author}/status/{id}` | /// Canonical URL: `x.com/{author}/status/{id}` |
@@ -32,8 +32,16 @@ pub struct Fetched {
pub source_url: String, pub source_url: String,
/// The exact HTML produced by the site's caption(). /// The exact HTML produced by the site's caption().
pub caption: String, pub caption: String,
/// Raw post text (tweet text / bsky text / pixiv title). /// The post's own title, where the platform has one: a pixiv artwork's
/// title, the headline of a bilibili opus post or the title of the video
/// an AV dynamic attaches. Empty on the platforms whose posts are text
/// only (x/twitter, bsky, misskey) and on bilibili posts without a
/// headline.
pub title: String, pub title: String,
/// The post's body text, as the platform exposes it: a tweet, a bsky or
/// misskey post, a bilibili dynamic's text, a pixiv artwork's description
/// (HTML flattened). Empty when the post has no text at all.
pub content: String,
pub media: Vec<crate::media::Media>, pub media: Vec<crate::media::Media>,
/// Spoiler flag for all media of this post. /// Spoiler flag for all media of this post.
pub sensitive: bool, pub sensitive: bool,
@@ -48,24 +56,38 @@ pub struct Fetched {
pub(crate) _keep_alive: Option<tempfile::TempDir>, pub(crate) _keep_alive: Option<tempfile::TempDir>,
} }
/// Values for the `{url} {author} {author_url} {title} {tags}` placeholders in /// Values for the `{url} {author} {author_url} {title} {content} {tags}`
/// user-supplied caption formats, substituted by [`caption_from_fields`] as /// placeholders in user-supplied caption formats, substituted by
/// HTML text (never as an attribute value). /// [`caption_from_fields`] as HTML text (never as an attribute value).
/// ///
/// `author`, `title` and `tags` come from the site API (post text, display /// `author`, `title`, `content` and `tags` come from the site API (post
/// names) and are HTML-escaped at construction. `url` and `author_url` stay /// text, display names, descriptions) and are HTML-escaped at construction.
/// raw: they are canonical URLs the adapter builds from numeric ids and /// `url` and `author_url` stay raw: they are canonical URLs the adapter
/// API-constrained handles/DIDs, so they carry no escapable character — the /// builds from numeric ids and API-constrained handles/DIDs, so they carry
/// bot's `/test` report relies on that when it embeds them. /// no escapable character — the bot's `/test` report relies on that when it
/// embeds them.
#[derive(Debug)] #[derive(Debug)]
pub(crate) struct RenderData { pub(crate) struct RenderData {
pub url: String, pub url: String,
pub author: String, pub author: String,
pub author_url: String, pub author_url: String,
pub title: String, pub title: String,
pub content: String,
pub tags: String, pub tags: String,
} }
/// The post's text as one string: title and content joined by a line break,
/// each only when it is non-empty. This is what the sites' built-in captions
/// show after the author line, and what the bot quotes when it is long.
pub fn compose_text(title: &str, content: &str) -> String {
match (title.is_empty(), content.is_empty()) {
(false, false) => format!("{title}\n{content}"),
(false, true) => title.to_string(),
(true, false) => content.to_string(),
(true, true) => String::new(),
}
}
impl Fetched { impl Fetched {
/// The site this post came from (used for per-site format overrides). /// The site this post came from (used for per-site format overrides).
/// A thin alias over [`Fetched::site_id`] kept for callers that read the /// A thin alias over [`Fetched::site_id`] kept for callers that read the
@@ -89,21 +111,23 @@ impl Fetched {
&data.author, &data.author,
&data.author_url, &data.author_url,
&data.title, &data.title,
&data.content,
&data.tags, &data.tags,
), ),
_ => truncate_caption(&self.caption), _ => truncate_caption(&self.caption),
} }
} }
/// The pre-escaped placeholder values (author, author_url, title, tags) /// The pre-escaped placeholder values (author, author_url, title,
/// a caller needs to rebuild a caption later, e.g. for a cached post /// content, tags) a caller needs to rebuild a caption later, e.g. for a
/// where the [`Fetched`] is no longer available. /// cached post where the [`Fetched`] is no longer available.
pub fn render_fields(&self) -> Option<(&str, &str, &str, &str)> { pub fn render_fields(&self) -> Option<(&str, &str, &str, &str, &str)> {
self.render_data.as_ref().map(|d| { self.render_data.as_ref().map(|d| {
( (
d.author.as_str(), d.author.as_str(),
d.author_url.as_str(), d.author_url.as_str(),
d.title.as_str(), d.title.as_str(),
d.content.as_str(),
d.tags.as_str(), d.tags.as_str(),
) )
}) })
@@ -149,6 +173,11 @@ pub fn truncate_caption(caption: &str) -> String {
/// [`Fetched::caption_with`]. An empty format returns `built_in` unchanged. /// [`Fetched::caption_with`]. An empty format returns `built_in` unchanged.
/// The result is truncated to [`MAX_CAPTION_CHARS`] (Telegram's caption /// The result is truncated to [`MAX_CAPTION_CHARS`] (Telegram's caption
/// limit for HTML parse mode). /// limit for HTML parse mode).
///
/// One flat argument per placeholder keeps the two callers (the fresh and the
/// cached caption path) mirroring each other; the same shape as the bot's
/// `debug_report`.
#[allow(clippy::too_many_arguments)]
pub fn caption_from_fields( pub fn caption_from_fields(
format: &str, format: &str,
built_in: &str, built_in: &str,
@@ -156,6 +185,7 @@ pub fn caption_from_fields(
author: &str, author: &str,
author_url: &str, author_url: &str,
title: &str, title: &str,
content: &str,
tags: &str, tags: &str,
) -> String { ) -> String {
if format.is_empty() { if format.is_empty() {
@@ -168,6 +198,7 @@ pub fn caption_from_fields(
.replace("{author}", author) .replace("{author}", author)
.replace("{author_url}", author_url) .replace("{author_url}", author_url)
.replace("{title}", title) .replace("{title}", title)
.replace("{content}", content)
.replace("{tags}", tags), .replace("{tags}", tags),
) )
} }
@@ -598,25 +629,34 @@ mod tests {
// The format string is escaped, the field values are substituted // The format string is escaped, the field values are substituted
// verbatim (callers pass the already-escaped render data). // verbatim (callers pass the already-escaped render data).
let out = caption_from_fields( let out = caption_from_fields(
"see {author} at {url} — {title}", "see {author} at {url} — {title}: {content}",
"", "",
"https://x.com/u/status/1", "https://x.com/u/status/1",
"A &amp; B", "A &amp; B",
"https://x.com/u", "https://x.com/u",
"hello <world>", "hello <world>",
"the body",
"", "",
); );
assert_eq!( assert_eq!(
out, out,
"see A &amp; B at https://x.com/u/status/1 — hello <world>" "see A &amp; B at https://x.com/u/status/1 — hello <world>: the body"
); );
// Empty format keeps the built-in caption untouched. // Empty format keeps the built-in caption untouched.
assert_eq!( assert_eq!(
caption_from_fields("", "built-in", "u", "a", "au", "t", "g"), caption_from_fields("", "built-in", "u", "a", "au", "t", "c", "g"),
"built-in" "built-in"
); );
} }
#[test]
fn compose_text_joins_title_and_content() {
assert_eq!(compose_text("标题", "正文"), "标题\n正文");
assert_eq!(compose_text("标题", ""), "标题");
assert_eq!(compose_text("", "正文"), "正文");
assert_eq!(compose_text("", ""), "");
}
#[test] #[test]
fn truncate_caption_keeps_short_text() { fn truncate_caption_keeps_short_text() {
assert_eq!(truncate_caption("short"), "short"); assert_eq!(truncate_caption("short"), "short");
@@ -107,10 +107,60 @@ pub fn media_headers(url: &str) -> Option<Vec<(&'static str, String)>> {
} }
} }
/// Flattens the app API's HTML description into plain text: `<br>` (and `<p>`)
/// become line breaks, other tags are dropped, entities decoded, the ends
/// trimmed. A caption shows text, not markup, so the author's `<a href>` links
/// contribute their link text only.
fn flatten_html(raw: &str) -> String {
let mut out = String::with_capacity(raw.len());
let mut chars = raw.chars().peekable();
while let Some(c) = chars.next() {
// Only `<` followed by `/` or a letter opens a tag — a bare `<` in
// prose ("2 < 3") is text.
let opens_tag = c == '<'
&& chars
.peek()
.is_some_and(|next| *next == '/' || next.is_ascii_alphabetic());
if !opens_tag {
out.push(c);
continue;
}
let mut tag = String::new();
let mut closed = false;
for c in chars.by_ref() {
if c == '>' {
closed = true;
break;
}
tag.push(c);
}
if !closed {
// Unclosed `<…`: keep it as text rather than dropping the tail.
out.push('<');
out.push_str(&tag);
break;
}
// `<br>`, `<br/>`, `<br />` with or without attributes, and both
// halves of a paragraph break the line; everything else is dropped.
let tag = tag
.trim()
.trim_start_matches('/')
.trim_end_matches('/')
.trim()
.to_ascii_lowercase();
if tag == "p" || tag.starts_with("br") {
out.push('\n');
}
}
html_escape::decode_html_entities(&out).trim().to_string()
}
#[derive(Debug)] #[derive(Debug)]
pub struct Illustration { pub struct Illustration {
id: String, id: String,
title: String, title: String,
/// The artwork's description, HTML flattened to plain text.
content: String,
author: String, author: String,
author_id: String, author_id: String,
tags: Vec<String>, tags: Vec<String>,
@@ -150,6 +200,7 @@ impl Illustration {
pub fn from_model(model: &IllustrationModel) -> Self { pub fn from_model(model: &IllustrationModel) -> Self {
let id = model.id.to_string(); let id = model.id.to_string();
let title = model.title.clone(); let title = model.title.clone();
let content = flatten_html(&model.caption);
let author = model.user.name.clone(); let author = model.user.name.clone();
let author_id = model.user.id.to_string(); let author_id = model.user.id.to_string();
let mut tags: Vec<String> = model.tags.iter().map(|tag| tag.name.clone()).collect(); let mut tags: Vec<String> = model.tags.iter().map(|tag| tag.name.clone()).collect();
@@ -193,6 +244,7 @@ impl Illustration {
Self { Self {
id, id,
title, title,
content,
author, author,
author_id, author_id,
tags, tags,
@@ -218,12 +270,14 @@ impl From<Illustration> for Fetched {
author: encode_text(&illustration.author).into_owned(), author: encode_text(&illustration.author).into_owned(),
author_url: author_url.clone(), author_url: author_url.clone(),
title: encode_text(&illustration.title).into_owned(), title: encode_text(&illustration.title).into_owned(),
content: encode_text(&illustration.content).into_owned(),
tags: encode_text(&tags).into_owned(), tags: encode_text(&tags).into_owned(),
}); });
Fetched { Fetched {
source_url: url, source_url: url,
caption: illustration.caption(), caption: illustration.caption(),
title: illustration.title.clone(), title: illustration.title.clone(),
content: illustration.content.clone(),
media: illustration.media, media: illustration.media,
sensitive: illustration.nsfw, sensitive: illustration.nsfw,
site_id: "pixiv", site_id: "pixiv",
@@ -262,6 +316,7 @@ mod tests {
"illust": { "illust": {
"id": 123, "id": 123,
"title": "Art <title>", "title": "Art <title>",
"caption": "一行说明<br />二行 <a href=\"https://x.example/\">链接</a> &amp; 结尾",
"type": type_, "type": type_,
"image_urls": { "image_urls": {
"medium": "medium.jpg", "medium": "medium.jpg",
@@ -284,6 +339,44 @@ mod tests {
Illustration::from_model(&model) Illustration::from_model(&model)
} }
/// The description arrives as HTML and becomes plain-text content: breaks
/// kept, tags dropped (links keep their text), entities decoded.
#[test]
fn from_json_maps_description_to_content() {
let v = illust_json("illust", 1, None, Some("o.jpg"), vec![], 0);
let illustration = parse(v);
assert_eq!(illustration.content, "一行说明\n二行 链接 & 结尾");
let fetched: Fetched = illustration.into();
assert_eq!(fetched.title, "Art <title>");
assert_eq!(fetched.content, "一行说明\n二行 链接 & 结尾");
// The built-in caption keeps its layout: the description stays out of
// it and is available through `{content}`.
assert!(!fetched.caption.contains("一行说明"), "{}", fetched.caption);
assert_eq!(
fetched.render_fields().unwrap().3,
"一行说明\n二行 链接 &amp; 结尾"
);
assert!(
fetched
.caption_with("{title}: {content}")
.ends_with("一行说明\n二行 链接 &amp; 结尾")
);
}
#[test]
fn flatten_html_handles_common_markup() {
assert_eq!(flatten_html(""), "");
assert_eq!(flatten_html("plain"), "plain");
assert_eq!(flatten_html("a<br />b<br/>c<br>d"), "a\nb\nc\nd");
// A paragraph break is a blank line, exactly like `<br /><br />` —
// writing it as one newline would flatten the author's paragraphs.
assert_eq!(flatten_html("<p>one</p><p>two</p>"), "one\n\ntwo");
assert_eq!(flatten_html("a &amp; b &lt;c&gt;"), "a & b <c>");
// Nothing to strip: angle brackets that are not a tag survive.
assert_eq!(flatten_html("2 < 3"), "2 < 3");
}
#[test] #[test]
fn pattern_matches_all_forms() { fn pattern_matches_all_forms() {
let cases = [ let cases = [
+4
View File
@@ -6,6 +6,10 @@ use serde::Deserialize;
pub struct IllustrationModel { pub struct IllustrationModel {
pub id: u64, pub id: u64,
pub title: String, pub title: String,
/// The artwork's description as the app API returns it — HTML in most
/// works (`<br />`, `<a href>`, sometimes `<p>`), empty for many.
#[serde(default)]
pub caption: String,
pub r#type: TypeModel, pub r#type: TypeModel,
pub image_urls: ImageUrlsModel, pub image_urls: ImageUrlsModel,
pub user: UserInfoModel, pub user: UserInfoModel,
+2 -1
View File
@@ -327,7 +327,8 @@ mod tests {
"https://x.com/nsfw_author/status/2083868672721039569" "https://x.com/nsfw_author/status/2083868672721039569"
); );
// The appended media short link (no URL-entity mapping) is stripped. // The appended media short link (no URL-entity mapping) is stripped.
assert_eq!(fetched.title, "nsfw content"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "nsfw content");
} }
#[test] #[test]
+10 -4
View File
@@ -99,6 +99,7 @@ fn empty_fetched(url: &str) -> Fetched {
// crafted links cannot break the parse (Telegram 400). // crafted links cannot break the parse (Telegram 400).
caption: encode_text(url).into_owned(), caption: encode_text(url).into_owned(),
title: String::new(), title: String::new(),
content: String::new(),
media: vec![], media: vec![],
sensitive: true, sensitive: true,
site_id: "twitter", site_id: "twitter",
@@ -352,17 +353,20 @@ impl From<Tweet> for Fetched {
fn from(tweet: Tweet) -> Self { fn from(tweet: Tweet) -> Self {
let url = tweet.url(); let url = tweet.url();
let author_url = tweet.author_url(); let author_url = tweet.author_url();
// A tweet has no title: its text is all content.
let render_data = Some(crate::site::RenderData { let render_data = Some(crate::site::RenderData {
url: url.clone(), url: url.clone(),
author: encode_text(&tweet.author).into_owned(), author: encode_text(&tweet.author).into_owned(),
author_url: author_url.clone(), author_url: author_url.clone(),
title: encode_text(&tweet.text).into_owned(), title: String::new(),
content: encode_text(&tweet.text).into_owned(),
tags: String::new(), tags: String::new(),
}); });
Fetched { Fetched {
source_url: url, source_url: url,
caption: tweet.caption(), caption: tweet.caption(),
title: tweet.text.clone(), title: String::new(),
content: tweet.text.clone(),
media: tweet.media, media: tweet.media,
sensitive: tweet.sensitive, sensitive: tweet.sensitive,
site_id: "twitter", site_id: "twitter",
@@ -437,7 +441,8 @@ mod tests {
assert_eq!(tweet.text, ">^ω^< & more 'quoted'"); assert_eq!(tweet.text, ">^ω^< & more 'quoted'");
assert_eq!(tweet.author, "O'Brien"); assert_eq!(tweet.author, "O'Brien");
let fetched: Fetched = tweet.into(); let fetched: Fetched = tweet.into();
assert_eq!(fetched.title, ">^ω^< & more 'quoted'"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, ">^ω^< & more 'quoted'");
// The caption escapes the raw text exactly once (encode_text covers // The caption escapes the raw text exactly once (encode_text covers
// & < >; apostrophes stay literal — they are harmless in text). // & < >; apostrophes stay literal — they are harmless in text).
assert!( assert!(
@@ -496,7 +501,8 @@ mod tests {
fetched.source_url, fetched.source_url,
"https://x.com/author_handle/status/861627479294746624" "https://x.com/author_handle/status/861627479294746624"
); );
assert_eq!(fetched.title, "a & b <c>"); assert_eq!(fetched.title, "");
assert_eq!(fetched.content, "a & b <c>");
assert!(fetched.sensitive); assert!(fetched.sensitive);
assert_eq!(fetched.media.len(), 2); assert_eq!(fetched.media.len(), 2);
match &fetched.media[0] { match &fetched.media[0] {
+18 -5
View File
@@ -419,6 +419,7 @@ pub(crate) async fn execute_command(
fetched.site_name(), fetched.site_name(),
&fetched.source_url, &fetched.source_url,
&fetched.title, &fetched.title,
&fetched.content,
fetched.render_fields(), fetched.render_fields(),
fetched.sensitive, fetched.sensitive,
&fetched.caption, &fetched.caption,
@@ -471,7 +472,8 @@ fn debug_report(
site_id: &str, site_id: &str,
source_url: &str, source_url: &str,
title: &str, title: &str,
render: Option<(&str, &str, &str, &str)>, content: &str,
render: Option<(&str, &str, &str, &str, &str)>,
sensitive: bool, sensitive: bool,
caption: &str, caption: &str,
media: &[x_media::media::Media], media: &[x_media::media::Media],
@@ -491,7 +493,8 @@ fn debug_report(
html_escape::encode_text(source_url) html_escape::encode_text(source_url)
)); ));
lines.push(format!("title: {}", html_escape::encode_text(title))); lines.push(format!("title: {}", html_escape::encode_text(title)));
if let Some((author, author_url, _title, tags)) = render { lines.push(format!("content: {}", html_escape::encode_text(content)));
if let Some((author, author_url, _title, _content, tags)) = render {
// The render fields are already pre-escaped for HTML captions; embed // The render fields are already pre-escaped for HTML captions; embed
// them as-is so the report renders them exactly like the final // them as-is so the report renders them exactly like the final
// caption. `author_url` is raw and gets escaped here. // caption. `author_url` is raw and gets escaped here.
@@ -559,7 +562,14 @@ mod tests {
"twitter", "twitter",
"https://x.com/u/status/1", "https://x.com/u/status/1",
"My title", "My title",
Some(("Author", "https://x.com/u", "My title", "tag1 tag2")), "My content",
Some((
"Author",
"https://x.com/u",
"My title",
"My content",
"tag1 tag2",
)),
false, false,
"<a href=\"https://x.com/u\">Author</a> · My title", "<a href=\"https://x.com/u\">Author</a> · My title",
&media, &media,
@@ -567,6 +577,7 @@ mod tests {
assert!(report.contains("site: twitter"), "{report}"); assert!(report.contains("site: twitter"), "{report}");
assert!(report.contains("key: twitter:1"), "{report}"); assert!(report.contains("key: twitter:1"), "{report}");
assert!(report.contains("title: My title"), "{report}"); assert!(report.contains("title: My title"), "{report}");
assert!(report.contains("content: My content"), "{report}");
assert!(report.contains("author: Author"), "{report}"); assert!(report.contains("author: Author"), "{report}");
assert!(report.contains("author_url: https://x.com/u"), "{report}"); assert!(report.contains("author_url: https://x.com/u"), "{report}");
assert!(report.contains("tags: tag1 tag2"), "{report}"); assert!(report.contains("tags: tag1 tag2"), "{report}");
@@ -584,7 +595,7 @@ mod tests {
#[test] #[test]
fn debug_report_without_render_data_and_no_media() { fn debug_report_without_render_data_and_no_media() {
let report = debug_report("u", "pixiv", "s", "t", None, true, "c", &[]); let report = debug_report("u", "pixiv", "s", "t", "c", None, true, "p", &[]);
assert!(!report.contains("author:"), "{report}"); assert!(!report.contains("author:"), "{report}");
assert!(report.contains("sensitive: true"), "{report}"); assert!(report.contains("sensitive: true"), "{report}");
assert!(report.contains("media (0):"), "{report}"); assert!(report.contains("media (0):"), "{report}");
@@ -601,10 +612,12 @@ mod tests {
"twitter", "twitter",
"https://x.com/u/status/1", "https://x.com/u/status/1",
"A & B <C>", "A & B <C>",
"body & <more>",
Some(( Some((
"A &amp; B", "A &amp; B",
"https://x.com/u", "https://x.com/u",
"A &amp; B &lt;C&gt;", "A &amp; B &lt;C&gt;",
"body &amp; &lt;more&gt;",
"#a &amp; #b", "#a &amp; #b",
)), )),
false, false,
@@ -640,7 +653,7 @@ mod tests {
fallback_url: None, fallback_url: None,
}) })
.collect(); .collect();
let report = debug_report("u", "twitter", "s", "t", None, false, "c", &media); let report = debug_report("u", "twitter", "s", "t", "c", None, false, "p", &media);
assert!(report.chars().count() <= MAX_DEBUG_REPORT_CHARS, "{report}"); assert!(report.chars().count() <= MAX_DEBUG_REPORT_CHARS, "{report}");
assert!(report.ends_with('…'), "{report}"); assert!(report.ends_with('…'), "{report}");
} }
+16 -12
View File
@@ -327,6 +327,7 @@ pub(crate) async fn url_media(
&cached.author, &cached.author,
&cached.author_url, &cached.author_url,
&cached.title, &cached.title,
&cached.content,
&cached.tags, &cached.tags,
) )
}; };
@@ -406,18 +407,20 @@ pub(crate) async fn url_media(
let caption = fetched.caption_with(&format); let caption = fetched.caption_with(&format);
// Raw render data for the link cache; the send fills in the // Raw render data for the link cache; the send fills in the
// Telegram file ids and persists the entry. // Telegram file ids and persists the entry.
let cache_data = fetched let cache_data =
.render_fields() fetched
.map(|(author, author_url, title, tags)| CachedPost { .render_fields()
url: fetched.source_url.clone(), .map(|(author, author_url, title, content, tags)| CachedPost {
caption: fetched.caption.clone(), url: fetched.source_url.clone(),
title: title.to_string(), caption: fetched.caption.clone(),
author: author.to_string(), title: title.to_string(),
author_url: author_url.to_string(), content: content.to_string(),
tags: tags.to_string(), author: author.to_string(),
sensitive: fetched.sensitive, author_url: author_url.to_string(),
media: vec![], tags: tags.to_string(),
}); sensitive: fetched.sensitive,
media: vec![],
});
let items: Vec<MediaItemPayload> = fetched let items: Vec<MediaItemPayload> = fetched
.media .media
.iter() .iter()
@@ -465,6 +468,7 @@ mod tests {
url: "https://x.com/u/status/1".into(), url: "https://x.com/u/status/1".into(),
caption: "cap".into(), caption: "cap".into(),
title: "t".into(), title: "t".into(),
content: "c".into(),
author: "a".into(), author: "a".into(),
author_url: "au".into(), author_url: "au".into(),
tags: "".into(), tags: "".into(),
+5
View File
@@ -38,6 +38,10 @@ pub struct CachedPost {
/// override). /// override).
pub caption: String, pub caption: String,
pub title: String, pub title: String,
/// The post's body text. Defaulted on read: entries written before the
/// title/content split carry it inside `title`.
#[serde(default)]
pub content: String,
pub author: String, pub author: String,
pub author_url: String, pub author_url: String,
pub tags: String, pub tags: String,
@@ -182,6 +186,7 @@ mod tests {
url: "https://x.com/u/status/1".into(), url: "https://x.com/u/status/1".into(),
caption: "cap".into(), caption: "cap".into(),
title: "t".into(), title: "t".into(),
content: "c".into(),
author: "a".into(), author: "a".into(),
author_url: "au".into(), author_url: "au".into(),
tags: "".into(), tags: "".into(),
+1
View File
@@ -1259,6 +1259,7 @@ mod tests {
url: "https://x.com/u/status/1".into(), url: "https://x.com/u/status/1".into(),
caption: "cap".into(), caption: "cap".into(),
title: "t".into(), title: "t".into(),
content: "c".into(),
author: "a".into(), author: "a".into(),
author_url: "au".into(), author_url: "au".into(),
tags: String::new(), tags: String::new(),
+1 -1
View File
@@ -18,7 +18,7 @@ pub struct ChatData {
/// name -> HTML template containing "[]" /// name -> HTML template containing "[]"
pub template: HashMap<String, String>, pub template: HashMap<String, String>,
/// site name (twitter/bsky/misskey/pixiv/bilibili) -> user-supplied caption format /// site name (twitter/bsky/misskey/pixiv/bilibili) -> user-supplied caption format
/// with {url} {author} {author_url} {title} {tags} placeholders. /// with {url} {author} {author_url} {title} {content} {tags} placeholders.
pub message_format: HashMap<String, String>, pub message_format: HashMap<String, String>,
} }