arXiv:2602.10675cs.CVcs.AI2026-02

构建首个大规模动态视觉推理数据集,提升视频问答中的时序推理能力。

TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning

  • 基于270万视频片段构建时序感知的视觉推理数据集
  • 在1078个样本上验证推理轨迹与答案正确性,性能超越现有方法
  • 提出统一模型生成未来帧与文本推理,支持动态场景问答

视觉思维链(VCoT)通过将视觉感知融入中间推理步骤,为多模态推理提供了新范式。然而,现有VCoT方法多局限于静态场景,难以捕捉指令、预测和摄像机运动等任务所需的时序动态。为此,我们提出TwiFF-2.7M,首个基于270万视频片段构建的大规模时序化VCoT数据集,专为动态视觉问答设计。同时,我们构建了包含1,078个样本的TwiFF-Bench高质量评估基准,用于评估开放动态环境下推理轨迹的合理性与最终答案的正确性。在此基础上,我们提出TwiFF模型,一种融合预训练视频生成与图像理解能力的统一架构,通过迭代生成未来动作帧与文本推理,实现时序一致的视觉推理。大量实验表明,TwiFF显著优于现有VCoT方法与文本思维链基线,在动态推理任务中充分验证了其有效性。代码与数据已公开于https://github.com/LiuJunhua02/TwiFF。

原文摘要 · Abstract (English)

Visual Chain-of-Thought (VCoT) has emerged as a promising paradigm for enhancing multimodal reasoning by integrating visual perception into intermediate reasoning steps. However, existing VCoT approaches are largely confined to static scenarios and struggle to capture the temporal dynamics essential for tasks such as instruction, prediction, and camera motion. To bridge this gap, we propose TwiFF-2.7M, the first large-scale, temporally grounded VCoT dataset derived from $2.7$ million video clips, explicitly designed for dynamic visual question and answer. Accompanying this, we introduce TwiFF-Bench, a high-quality evaluation benchmark of $1,078$ samples that assesses both the plausibility of reasoning trajectories and the correctness of final answers in open-ended dynamic settings. Building on these foundations, we propose the TwiFF model, a unified modal that synergistically leverages pre-trained video generation and image comprehension capabilities to produce temporally coherent visual reasoning cues-iteratively generating future action frames and textual reasoning. Extensive experiments demonstrate that TwiFF significantly outperforms existing VCoT methods and Textual Chain-of-Thought baselines on dynamic reasoning tasks, which fully validates the effectiveness for visual question answering in dynamic scenarios. Our code and data is available at https://github.com/LiuJunhua02/TwiFF.

视觉推理视频问答时序建模思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。