让视频检索理解故事背后的动机,突破传统动作识别局限。
StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning

- 构建需心智理论推理的叙事视频检索基准StoryTR
- 模型在复杂叙事线索中定位准确率仅0.53平均重叠度
- 用三层心智链数据训练的7B模型提升15.1%相对精度
当前视频时刻检索擅长动作识别,却难以理解叙事内容。模型能看见‘发生了什么’,却无法推断‘为何重要’。这一语义鸿沟源于缺乏‘心智理论(ToM)’——即从表层观察推断隐含意图、心理状态与叙事因果的能力。本文提出首个需ToM推理的视频时刻检索基准StoryTR,包含8.1k条来自叙事性短视频(shorts/reels)的样本。这些视频信息密度高,意义通过细微多模态线索传递,如一个眼神加叹息的组合含义远超眼神本身。单纯多模态感知不足,需结合心智理论解码‘微笑’可能隐藏‘敌意’。为此,我们设计了‘代理式数据流水线’,生成包含三阶段ToM链(意图解析、叙事推理、边界定位)的训练数据。实验显示推理差距严重:Gemini-3.0-Pro在StoryTR上仅获0.53平均交并比。而基于该数据训练的7B Shorts-Moment模型,相比基线提升15.1%相对交并比,证明‘叙事推理能力’比参数规模更重要。
原文摘要 · Abstract (English)
Current video moment retrieval excels at action-centric tasks but struggles with narrative content. Models can see \textit{what is happening} but fail to reason \textit{why it matters}. This semantic gap stems from the lack of \textbf{Theory of Mind (ToM)}: the cognitive ability to infer implicit intentions, mental states, and narrative causality from surface-level observations. We introduce \textbf{StoryTR}, the first video moment retrieval benchmark requiring ToM reasoning, comprising 8.1k samples from narrative short-form videos (shorts/reels). These videos present an ideal testbed. Their high information density encodes meaning through subtle multimodal cues. For instance, a glance paired with a sigh carries entirely different semantics than the glance alone. Yet multimodal perception alone is insufficient; ToM is required to decode that a character ``smiling'' may actually be ``concealing hostility.'' To teach models this reasoning capability, we propose an \textbf{Agentic Data Pipeline} that generates training data with explicit three-tier ToM chains (intent decoding, narrative reasoning, boundary localization). Experiments reveal the severity of the reasoning gap: Gemini-3.0-Pro achieves only 0.53 Avg IoU on StoryTR. However, our 7B \textbf{Shorts-Moment} model, trained on ToM-guided data, improves +15.1\% relative IoU over baselines, demonstrating that \textit{narrative reasoning capability matters more than parameter scale}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。