Compare commits

...
6 Commits
11 changed files with 457 additions and 23 deletions
+3 -3
View File
@@ -27,7 +27,7 @@ The `x-media` library: `site::fetch(url)` dispatches (in order) twitter → bsky
| Path | Purpose |
|---|---|
| `crates/x-media/src/` | Fetch library. `site/mod.rs` = dispatcher + `Fetched`/`FetchError`/`download_media`/`media_size`; `media.rs` = `Media` enum; `examples/fetch.rs` = end-to-end usage sample |
| `crates/x-media/src/site/<twitter\|pixiv\|bsky>/` | One directory per site: `mod.rs` (re-exports), `interface.rs` (PATTERN, `enabled()`, `fetch_from_url()`, site struct, `From<SiteStruct> for Fetched`), `model.rs` (serde DTOs). Pixiv adds `api.rs` (auth + transport) |
| `crates/x-media/src/site/<twitter\|pixiv\|bsky>/` | One directory per site: `mod.rs` (re-exports), `interface.rs` (PATTERN, `enabled()`, `fetch_from_url()`, site struct, `From<SiteStruct> for Fetched`), `model.rs` (serde DTOs). Pixiv adds `api.rs` (auth + transport); twitter adds `auth.rs` (logged-in GraphQL `TweetDetail` fallback for NSFW tweets, gated on `TWITTER_AUTH_TOKEN`) |
| `crates/xmedia-bot/src/main.rs` | Entry point: env/log init, queue worker start, pixiv validation, 300 s edit-expiry sweep, dptree handler tree, webhook vs polling dispatch |
| `crates/xmedia-bot/src/config.rs` | Manual env parsing into `Config` |
| `crates/xmedia-bot/src/handlers.rs` | `Command` enum (teloxide `BotCommands`), message/inline/callback handlers, URL extraction, global statics |
@@ -72,7 +72,7 @@ Docker: `docker build -t tgxmb .` then `docker run --rm -d --name tgxmb --env-fi
| `crates/x-media/src/site/pixiv/api.rs` | OAuth token exchange (hardcoded app client id/secret), access-token cache, ugoira zip→MP4 via ffmpeg in `spawn_blocking` |
| `Dockerfile` | Multi-stage: cached dep layer via stub sources + `touch *.rs` mtime hack, static ffmpeg from ffmpeg.martin-riedl.de (`FFMPEG_URL` arg, `unzip -t` integrity check), `debian:bookworm-slim` runtime, entrypoint |
| `docker-entrypoint.sh` | Privilege drop: `useradd` with `LOCAL_USER_ID` (default 9001) + `setpriv` (no gosu on bookworm-slim) |
| `docker-compose.yml.example` | Deployment env reference (real `docker-compose.yml` is gitignored). Ships nginx-proxy + acme-companion: webhook mode needs TLS termination in front (teloxide's axum listener is HTTP-only; `WEBHOOK_CERT` only feeds `set_webhook`), bot exposes `VIRTUAL_HOST`/`VIRTUAL_PORT` on the shared `proxy` network, no host port |
| `docker-compose.yml.example` | Deployment env reference (real `docker-compose.yml` is gitignored). Ships nginx-proxy + acme-companion: webhook mode needs TLS termination in front (teloxide's axum listener is HTTP-only; `WEBHOOK_CERT` only feeds `set_webhook`), bot exposes `VIRTUAL_HOST`/`VIRTUAL_PORT` on the shared `proxy` network, no host port; container names `nginx-proxy`/`acme-companion`/`tgxmb`, start order via `depends_on` (proxy → acme → bot) |
| `.github/workflows/docker.yml` | CI: build+push to Docker Hub on tag `v*`/master; **no test step**; buildx gha cache (`cache-from`/`cache-to`, scope `tgxmb-build`, `mode=max`) so cargo deps + ffmpeg layers are restored across runs |
| `README.md` | Feature docs + command table (Chinese) |
@@ -81,7 +81,7 @@ Docker: `docker build -t tgxmb .` then `docker run --rm -d --name tgxmb --env-fi
- **Rust, stable, edition 2024**, workspace resolver 3. No `rust-version`/MSRV pin, no `rust-toolchain.toml` — recent stable is assumed. No nightly features.
- Package manager: **Cargo** (workspace with path dep `x-media``xmedia-bot`). No `[workspace.package]`/shared deps — each crate lists deps independently.
- **Two reqwest versions coexist in the lock** (0.12.28 via teloxide, 0.13.3 in x-media) — don't unify casually.
- Config is **environment-variable driven** (dotenv loads `.env`, gitignored; no `.env.example` exists). Key vars: `TELOXIDE_TOKEN` (required), `PIXIV_REFRESH_TOKEN`, `BOT_ADMIN` (comma-separated ids), `EDIT_MESSAGE_TTL_SECONDS` (default 86400), `WEBHOOK`/`WEBHOOK_URL`/`WEBHOOK_LISTEN`/`WEBHOOK_PORT`/`WEBHOOK_CERT`/`WEBHOOK_SECRET_TOKEN` (webhook mode requires URL/listen/port, `.expect`ed; `WEBHOOK_CERT` is Telegram-facing self-signed validation only — TLS must be terminated by a reverse proxy), `RUST_LOG`, `TELOXIDE_PROXY`, `LOCAL_USER_ID` (entrypoint only).
- Config is **environment-variable driven** (dotenv loads `.env`, gitignored; no `.env.example` exists). Key vars: `TELOXIDE_TOKEN` (required), `PIXIV_REFRESH_TOKEN`, `TWITTER_AUTH_TOKEN` (optional; x.com `auth_token` cookie — enables the logged-in GraphQL fallback that fetches NSFW tweets syndication withholds), `BOT_ADMIN` (comma-separated ids), `EDIT_MESSAGE_TTL_SECONDS` (default 86400), `WEBHOOK`/`WEBHOOK_URL`/`WEBHOOK_LISTEN`/`WEBHOOK_PORT`/`WEBHOOK_CERT`/`WEBHOOK_SECRET_TOKEN` (webhook mode requires URL/listen/port, `.expect`ed; `WEBHOOK_CERT` is Telegram-facing self-signed validation only — TLS must be terminated by a reverse proxy), `RUST_LOG`, `TELOXIDE_PROXY`, `LOCAL_USER_ID` (entrypoint only).
- SQLite via `rusqlite` with `bundled` feature (no system libsqlite needed). DB file `data/task_queue.db` is CWD-relative — run from the workspace root, or `/app` in Docker. Mount `./data` and `./cert` volumes.
- `.gitattributes` enforces LF for `*.sh` (CRLF breaks shebangs in containers). `.gitignore`: `.env`, `data/`, `cert/`, `docker-compose.yml`, `/target`, `.idea/`.
- Docs are in Chinese; user-facing bot strings too. Keep that convention when editing captions/templates/docs.
Generated
+3 -2
View File
@@ -3296,12 +3296,13 @@ checksum = "1ffae5123b2d3fc086436f8834ae3ab053a283cfac8fe0a0b8eaae044768a4c4"
[[package]]
name = "x-media"
version = "1.0.4"
version = "1.0.5"
dependencies = [
"bytes",
"dotenv",
"html-escape",
"log",
"rand 0.8.6",
"regex",
"reqwest 0.13.3",
"serde",
@@ -3314,7 +3315,7 @@ dependencies = [
[[package]]
name = "xmedia-bot"
version = "1.0.4"
version = "1.0.5"
dependencies = [
"dotenv",
"html-escape",
+3 -1
View File
@@ -28,7 +28,9 @@ docker build -t tgxmb .
docker run --rm -d --name tgxmb --env-file .env -v ./data:/app/data tgxmb
```
环境变量:`TELOXIDE_TOKEN`(必填)、`PIXIV_REFRESH_TOKEN``BOT_ADMIN``EDIT_MESSAGE_TTL_SECONDS``RUST_LOG``WEBHOOK*`
环境变量:`TELOXIDE_TOKEN`(必填)、`PIXIV_REFRESH_TOKEN``BOT_ADMIN``EDIT_MESSAGE_TTL_SECONDS``RUST_LOG``WEBHOOK*``TWITTER_AUTH_TOKEN`(可选)
NSFW 推文:公开的 syndication 接口不返回敏感内容。设置 `TWITTER_AUTH_TOKEN`(登录 x.com 后浏览器 Cookie 里的 `auth_token` 值)后,bot 会仅在遇到 NSFW 推文时以登录态获取媒体;未设置则提示无媒体。
### Webhook 部署(需要反向代理)
+2 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "x-media"
version = "1.0.4"
version = "1.0.5"
edition = "2024"
[dependencies]
@@ -13,6 +13,7 @@ url = "2.5.2"
bytes = "1"
zip = "2"
tempfile = "3"
rand = "0.8"
log = "0.4"
tokio = { version = "1.40", features = ["time"] }
+5 -1
View File
@@ -89,6 +89,9 @@ pub enum FetchError {
Pixiv(PixivError),
NotFound,
Blocked,
/// The post exists but its content is withheld (twitter NSFW /
/// age-restricted tweets come back as an empty `{}` from syndication).
Sensitive,
}
impl fmt::Display for FetchError {
@@ -99,6 +102,7 @@ impl fmt::Display for FetchError {
FetchError::Pixiv(e) => write!(f, "pixiv error: {e}"),
FetchError::NotFound => write!(f, "not found"),
FetchError::Blocked => write!(f, "blocked"),
FetchError::Sensitive => write!(f, "content withheld (sensitive)"),
}
}
}
@@ -109,7 +113,7 @@ impl std::error::Error for FetchError {
FetchError::Http(e) => Some(e),
FetchError::Json(e) => Some(e),
FetchError::Pixiv(e) => Some(e),
FetchError::NotFound | FetchError::Blocked => None,
FetchError::NotFound | FetchError::Blocked | FetchError::Sensitive => None,
}
}
}
+376
View File
@@ -0,0 +1,376 @@
//! Authenticated fallback for tweets the public syndication endpoint refuses
//! to serve (NSFW / age-restricted tweets come back as an empty `{}`).
//!
//! Mirrors nazurin's web API client ([`web.py`]) and is used *only* when
//! syndication reports [`FetchError::Sensitive`]: the private GraphQL
//! `TweetDetail` endpoint, authenticated with a browser session cookie from
//! `TWITTER_AUTH_TOKEN` (the `auth_token` cookie value of a logged-in x.com
//! session). A fresh random `ct0` is generated per call; X checks that the
//! `x-csrf-token` header matches the cookie, not that it issued the value.
//!
//! [`web.py`]: https://github.com/y-young/nazurin/blob/master/nazurin/sites/twitter/api/web.py
//!
//! # Caveats
//! - X rotates the GraphQL query id when it rolls the web app; if requests
//! start failing, update [`TWEET_DETAIL_QUERY_ID`]. Fresh references from
//! the actively maintained FxEmbed/FxEmbed: TweetDetail
//! `R9IzzyzQBV87-DOWpcvDmw`, TweetResultByRestId `f2sagi1jweVHFkTUIHzmMQ`
//! (the latter is anonymous and surfaces NSFW tweets as
//! `reason: NsfwLoggedOut`).
//! - `x-client-transaction-id` is only required for `SearchTimeline`
//! (verified against FxEmbed's `proxy/allowlist.ts`) — TweetDetail works
//! without it; no need for the nazurin home-page/JS-bundle derivation.
use std::sync::LazyLock;
use serde_json::{json, Value};
use crate::site::FetchError;
use super::interface::Tweet;
/// `auth_token` cookie of a logged-in x.com session; enables the fallback.
/// Trimmed: a CRLF `.env` (Windows) leaves a trailing `\r` on the value,
/// which would make the Cookie header invalid.
static AUTH_TOKEN: LazyLock<Option<String>> = LazyLock::new(|| {
std::env::var("TWITTER_AUTH_TOKEN")
.ok()
.map(|s| s.trim().to_string())
.filter(|s| !s.is_empty())
});
/// Public "logged in" client token used by the x.com web app.
const LOGGED_IN_BEARER: &str =
"Bearer AAAAAAAAAAAAAAAAAAAAANRILgAAAAAAnNwIzUejRCOuH5E6I8xnZz4puTs%3D1Zv7ttfk8LF81IUq16cHjhLTvJu4FA33AGWWjCpTnA";
/// `TweetDetail` query id (from nazurin; still valid as of 2026-08,
/// corroborated by the current FxEmbed build — see module caveats).
const TWEET_DETAIL_QUERY_ID: &str = "_8aYOgEDz35BrBcBal1-_w";
fn variables(id: &str) -> Value {
json!({
"focalTweetId": id,
"with_rux_injections": false,
"includePromotedContent": false,
"withCommunity": true,
"withQuickPromoteEligibilityTweetFields": false,
"withBirdwatchNotes": false,
"withVoice": true,
})
}
fn features() -> Value {
json!({
"rweb_video_screen_enabled": false,
"profile_label_improvements_pcf_label_in_post_enabled": true,
"rweb_tipjar_consumption_enabled": true,
"verified_phone_label_enabled": false,
"creator_subscriptions_tweet_preview_api_enabled": true,
"responsive_web_graphql_timeline_navigation_enabled": true,
"responsive_web_graphql_skip_user_profile_image_extensions_enabled": false,
"premium_content_api_read_enabled": false,
"communities_web_enable_tweet_community_results_fetch": true,
"c9s_tweet_anatomy_moderator_badge_enabled": true,
"responsive_web_grok_analyze_button_fetch_trends_enabled": false,
"responsive_web_grok_analyze_post_followups_enabled": true,
"responsive_web_jetfuel_frame": false,
"responsive_web_grok_share_attachment_enabled": true,
"articles_preview_enabled": true,
"responsive_web_edit_tweet_api_enabled": true,
"graphql_is_translatable_rweb_tweet_is_translatable_enabled": true,
"view_counts_everywhere_api_enabled": true,
"longform_notetweets_consumption_enabled": true,
"responsive_web_twitter_article_tweet_consumption_enabled": true,
"tweet_awards_web_tipping_enabled": false,
"responsive_web_grok_show_grok_translated_post": false,
"responsive_web_grok_analysis_button_from_backend": true,
"creator_subscriptions_quote_tweet_preview_enabled": false,
"freedom_of_speech_not_reach_fetch_enabled": true,
"standardized_nudges_misinfo": true,
"tweet_with_visibility_results_prefer_gql_limited_actions_policy_enabled": true,
"longform_notetweets_rich_text_read_enabled": true,
"longform_notetweets_inline_media_enabled": true,
"responsive_web_grok_image_annotation_enabled": true,
"responsive_web_enhance_cards_enabled": false,
})
}
/// Whether the authenticated fallback is available.
pub fn enabled() -> bool {
AUTH_TOKEN.is_some()
}
/// Fetches a tweet as the logged-in user via the private GraphQL API.
/// Returns the syndication-shaped [`Tweet`] (media included for NSFW posts).
pub async fn fetch(id: &str) -> Result<Tweet, FetchError> {
let token = AUTH_TOKEN
.as_deref()
.ok_or(FetchError::Sensitive)?;
// 16 random bytes as 32 hex chars: X rejects ct0 values of any other
// length with 403 code 353 ("matching csrf cookie and header").
let ct0: String = (0..16)
.map(|_| format!("{:02x}", rand::random::<u8>()))
.collect();
let response = crate::site::CLIENT
.get(format!(
"https://x.com/i/api/graphql/{TWEET_DETAIL_QUERY_ID}/TweetDetail"
))
.query(&[
("variables", variables(id).to_string()),
("features", features().to_string()),
])
.header("authorization", LOGGED_IN_BEARER)
.header("x-csrf-token", &ct0)
.header("x-twitter-auth-type", "OAuth2Session")
.header("cookie", format!("auth_token={token}; ct0={ct0}"))
.header("x-twitter-client-language", "en")
.header("x-twitter-active-user", "yes")
.header("referer", "https://x.com/")
.send()
.await?;
if !response.status().is_success() {
log::warn!("twitter auth fetch {id}: HTTP {}", response.status());
return Err(FetchError::NotFound);
}
let text = response.text().await?;
let json: Value = serde_json::from_str(&text)?;
let result = parse_tweet_result(&json, id)?;
let syndication_shape = to_syndication_shape(&result)
.ok_or_else(|| FetchError::Json(serde_json::Error::io(std::io::Error::new(
std::io::ErrorKind::InvalidData,
"missing tweet fields in GraphQL response",
))))?;
Tweet::from_syndication_json(&syndication_shape.to_string()).map_err(FetchError::Json)
}
/// Locates the tweet for `id` in a `TweetDetail` response and unwraps
/// visibility wrappers / retweets, mirroring nazurin's `_process_response`.
fn parse_tweet_result(json: &Value, id: &str) -> Result<Value, FetchError> {
if let Some(errors) = json.get("errors").and_then(|e| e.as_array()) {
let messages: Vec<&str> = errors
.iter()
.filter_map(|e| e.get("message").and_then(|m| m.as_str()))
.collect();
log::warn!("twitter auth fetch {id} failed: {}", messages.join("; "));
return Err(FetchError::NotFound);
}
let instructions = json
.pointer("/data/threaded_conversation_with_injections_v2/instructions")
.and_then(|v| v.as_array())
.ok_or(FetchError::NotFound)?;
for instruction in instructions {
if instruction.get("type").and_then(|t| t.as_str()) != Some("TimelineAddEntries") {
continue;
}
let entries = instruction
.get("entries")
.and_then(|e| e.as_array())
.ok_or(FetchError::NotFound)?;
let wanted = format!("tweet-{id}");
for entry in entries {
if entry.get("entryId").and_then(|i| i.as_str()) == Some(wanted.as_str()) {
let result = entry
.pointer("/content/itemContent/tweet_results/result")
.ok_or(FetchError::NotFound)?;
return normalize_tweet_result(result);
}
}
}
Err(FetchError::NotFound)
}
/// Unwraps TweetTombstone/TweetUnavailable errors, the
/// TweetWithVisibilityResults wrapper and retweets, returning the
/// `{core, legacy, ...}` tweet object.
fn normalize_tweet_result(result: &Value) -> Result<Value, FetchError> {
match result.get("__typename").and_then(|t| t.as_str()) {
Some("TweetTombstone") => {
let text = result
.pointer("/tombstone/text/text")
.and_then(|t| t.as_str())
.unwrap_or("tweet is unavailable");
log::warn!("twitter auth fetch: tombstone: {text}");
return Err(FetchError::NotFound);
}
Some("TweetUnavailable") => {
let reason = result
.get("reason")
.and_then(|r| r.as_str())
.unwrap_or("unknown");
log::warn!("twitter auth fetch: tweet unavailable: {reason}");
return Err(FetchError::NotFound);
}
_ => {}
}
// TweetWithVisibilityResults (e.g. limited replies) nests the real tweet.
let tweet = result.get("tweet").unwrap_or(result);
// A retweet's media lives on the original tweet.
if let Some(original) = tweet.pointer("/legacy/retweeted_status_result/result") {
return Ok(original.clone());
}
Ok(tweet.clone())
}
/// Maps a GraphQL `{core, legacy, ...}` tweet onto the syndication JSON
/// shape [`Tweet::from_syndication_json`] parses, so the existing text /
/// media handling (t.co expansion, `name=orig`, mp4 variant) is reused.
fn to_syndication_shape(tweet: &Value) -> Option<Value> {
let legacy = tweet.get("legacy")?;
let user = tweet.pointer("/core/user_results/result/legacy")?;
Some(json!({
"id_str": legacy.get("id_str"),
"text": legacy.get("full_text"),
"user": {
"name": user.get("name"),
"screen_name": user.get("screen_name"),
},
"possibly_sensitive": legacy.get("possibly_sensitive"),
"display_text_range": legacy.get("display_text_range"),
"entities": legacy.get("entities"),
"mediaDetails": legacy.pointer("/extended_entities/media"),
}))
}
#[cfg(test)]
mod tests {
use super::*;
fn tweet_result() -> Value {
json!({
"__typename": "Tweet",
"core": {
"user_results": {
"result": {
"legacy": { "name": "Display Name", "screen_name": "nsfw_author" }
}
}
},
"legacy": {
"id_str": "2083868672721039569",
"full_text": "nsfw content https://t.co/abc123",
"display_text_range": [0, 12],
"possibly_sensitive": true,
"entities": {
"urls": [
{ "url": "https://t.co/abc123", "expanded_url": "https://example.com/x" }
]
},
"extended_entities": {
"media": [
{
"type": "photo",
"media_url_https": "https://pbs.twimg.com/media/nsfw.jpg",
"original_info": { "width": 1200, "height": 800 }
},
{
"type": "video",
"media_url_https": "https://pbs.twimg.com/thumb.jpg",
"video_info": {
"variants": [
{ "content_type": "application/x-mpegURL", "url": "https://x.com/pl.m3u8" },
{ "content_type": "video/mp4", "url": "https://video.twimg.com/nsfw.mp4" }
]
}
}
]
}
}
})
}
fn conversation(tweet: Value) -> Value {
json!({
"data": {
"threaded_conversation_with_injections_v2": {
"instructions": [
{ "type": "TimelineAddEntries", "entries": [
{ "entryId": "tweet-2083868672721039569",
"content": { "itemContent": { "tweet_results": { "result": tweet } } } }
]}
]
}
}
})
}
#[test]
fn parses_graphql_tweet_into_fetched() {
let json = conversation(tweet_result());
let result = parse_tweet_result(&json, "2083868672721039569").unwrap();
let shape = to_syndication_shape(&result).unwrap();
let tweet = Tweet::from_syndication_json(&shape.to_string()).unwrap();
let fetched: crate::site::Fetched = tweet.into();
assert!(fetched.sensitive);
assert_eq!(fetched.media.len(), 2);
match &fetched.media[0] {
crate::media::Media::Illustration { url, .. } => {
assert_eq!(url, "https://pbs.twimg.com/media/nsfw.jpg?name=orig");
}
other => panic!("expected illustration, got {other:?}"),
}
match &fetched.media[1] {
crate::media::Media::Video { url, .. } => {
assert_eq!(url, "https://video.twimg.com/nsfw.mp4");
}
other => panic!("expected video, got {other:?}"),
}
assert_eq!(fetched.source_url, "https://x.com/nsfw_author/status/2083868672721039569");
// display_text_range cuts the trailing t.co link.
assert_eq!(fetched.title, "nsfw content");
}
#[test]
fn unwraps_retweet_to_original() {
let original = tweet_result();
let mut rt = tweet_result();
rt["legacy"]["retweeted_status_result"] = json!({ "result": original });
let json = conversation(rt);
let result = parse_tweet_result(&json, "2083868672721039569").unwrap();
assert!(result.pointer("/legacy/retweeted_status_result").is_none());
assert_eq!(result.pointer("/legacy/id_str").unwrap(), "2083868672721039569");
}
#[test]
fn error_response_maps_to_not_found() {
let json = json!({ "errors": [{ "message": "NsfwLoggedOut" }] });
assert!(matches!(
parse_tweet_result(&json, "1"),
Err(FetchError::NotFound)
));
}
#[test]
fn missing_entry_maps_to_not_found() {
let json = conversation(json!({ "__typename": "Tweet" }));
assert!(matches!(
parse_tweet_result(&json, "999"),
Err(FetchError::NotFound)
));
}
#[test]
fn tombstone_maps_to_not_found() {
let tombstone = json!({
"__typename": "TweetTombstone",
"tombstone": { "text": { "text": "Age-restricted adult content" } }
});
let json = conversation(tombstone);
assert!(matches!(
parse_tweet_result(&json, "2083868672721039569"),
Err(FetchError::NotFound)
));
}
#[test]
fn visibility_wrapper_unwraps() {
let inner = tweet_result();
let wrapped = json!({ "__typename": "TweetWithVisibilityResults", "tweet": inner });
let json = conversation(wrapped);
let result = parse_tweet_result(&json, "2083868672721039569").unwrap();
assert_eq!(result.get("__typename").unwrap(), "Tweet");
}
}
+55 -13
View File
@@ -19,7 +19,44 @@ pub async fn fetch_from_url(url: &str) -> Result<Fetched, FetchError> {
.and_then(|caps| caps.get(1))
.map(|m| m.as_str())
.ok_or(FetchError::NotFound)?;
Ok(fetch(id).await?.into())
match fetch(id).await {
Ok(tweet) => Ok(tweet.into()),
// Syndication withholds NSFW/age-restricted tweets (empty `{}`).
// Retry as the logged-in user when TWITTER_AUTH_TOKEN is set;
// otherwise degrade to an empty result (the bot replies
// "No media found").
Err(FetchError::Sensitive) => {
if super::auth::enabled() {
match super::auth::fetch(id).await {
Ok(tweet) => Ok(tweet.into()),
Err(e) => {
log::warn!("twitter auth fallback failed for {id}: {e}");
Ok(empty_fetched(url))
}
}
} else {
log::info!(
"tweet {id} is sensitive; set TWITTER_AUTH_TOKEN to fetch NSFW media"
);
Ok(empty_fetched(url))
}
}
Err(e) => Err(e),
}
}
/// A Fetched with no media for withheld tweets: the bot replies
/// "No media found" and moves on instead of erroring.
fn empty_fetched(url: &str) -> Fetched {
Fetched {
source_url: url.to_string(),
caption: url.to_string(),
title: String::new(),
media: vec![],
sensitive: true,
render_data: None,
_keep_alive: None,
}
}
/// Fetches a tweet from the syndication endpoint. Deleted/blocked tweets
@@ -44,6 +81,15 @@ pub async fn fetch(id: &str) -> Result<Tweet, FetchError> {
{
return Err(FetchError::NotFound);
}
// NSFW / age-restricted tweets exist but are served as an empty `{}` —
// they surface as FetchError::Sensitive so the caller can retry as a
// logged-in user.
if serde_json::from_str::<serde_json::Value>(&text)
.map(|v| v.get("id_str").is_none())
.unwrap_or(false)
{
return Err(FetchError::Sensitive);
}
Ok(Tweet::from_syndication_json(&text).map_err(FetchError::Json)?)
}
@@ -159,21 +205,17 @@ impl Tweet {
}
/// The raw syndication `text` ends with the appended media short link
/// (" https://t.co/wmI8McgXul"). `display_text_range` (UTF-16 indices) marks
/// the visible text; a regex strips any remaining trailing t.co link when the
/// range is absent or a tweet ends in a URL short link.
/// (" https://t.co/wmI8McgXul"). `display_text_range` marks the visible text;
/// a regex strips any remaining trailing t.co link when the range is absent
/// or a tweet ends in a URL short link.
///
/// X reports these indices in Unicode **code points**, not UTF-16 units
/// (verified against GraphQL responses containing emoji: cutting an emoji
/// tweet by UTF-16 units silently drops the character after the emoji).
fn strip_trailing_short_links(text: &str, display_text_range: Option<[usize; 2]>) -> String {
let mut out = match display_text_range {
Some([start, end]) if start < end => {
let units: Vec<u16> = text
.encode_utf16()
.skip(start)
.take(end - start)
.collect();
// Drop the replacement char that a surrogate cut at the boundary
// would produce (the range end is a valid UTF-16 boundary in
// practice, so this is just a safety net).
String::from_utf16_lossy(&units).replace('\u{FFFD}', "")
text.chars().skip(start).take(end - start).collect()
}
_ => text.to_string(),
};
+1
View File
@@ -1,3 +1,4 @@
mod auth;
mod interface;
mod model;
+1 -1
View File
@@ -10,7 +10,7 @@ pub struct SyndicationTweet {
#[serde(default)]
pub possibly_sensitive: Option<bool>,
/// Visible-text span; the raw `text` field has the appended media short
/// link after it. Indices are UTF-16 code units.
/// link after it. Indices are Unicode code points (not UTF-16 units).
#[serde(default, rename = "display_text_range")]
pub display_text_range: Option<[usize; 2]>,
#[serde(default)]
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "xmedia-bot"
version = "1.0.4"
version = "1.0.5"
edition = "2024"
[dependencies]
+7
View File
@@ -16,6 +16,7 @@ services:
networks: [proxy]
labels:
- 'com.github.jrcs.letsencrypt_nginx_proxy_companion.nginx_proxy=true'
container_name: nginx-proxy
acme-companion:
image: nginxproxy/acme-companion
@@ -29,6 +30,7 @@ services:
- ./nginx-html:/usr/share/nginx/html:rw
- ./nginx-acme:/etc/acme.sh
networks: [proxy]
container_name: acme-companion
depends_on:
- nginx-proxy
@@ -40,6 +42,9 @@ services:
TELOXIDE_TOKEN: ''
BOT_ADMIN: ''
PIXIV_REFRESH_TOKEN: ''
# Optional: x.com session cookie (auth_token) — fetches NSFW tweets
# that the public syndication endpoint withholds.
TWITTER_AUTH_TOKEN: ''
EDIT_MESSAGE_TTL_SECONDS: '86400'
RUST_LOG: 'info'
VIRTUAL_HOST: 'bot.example.com'
@@ -54,6 +59,8 @@ services:
volumes:
- ./data:/app/data
networks: [proxy]
depends_on:
- nginx-proxy
container_name: tgxmb
networks: