arXiv:2606.05736cs.CV2026-06被引 1

让视频推理同时看图说话,提升准确率和训练效率

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning

论文配图:VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning
图 1 · 摘自论文原文
  • 图文交替推理:在思考过程中动态结合视觉帧与文本步骤
  • 性能领先:同参数量模型中表现最优,训练速度更快
  • 自动生成高质量多模态推理数据,解决标注稀缺问题

视频推理旨在理解视频中的复杂时序事件与因果关系。近期,思维链(CoT)被引入该领域以提升推理准确性。然而,现有基于CoT的视频推理方法主要依赖纯文本信息进行逻辑推断,忽略了推理过程中关键的视觉信息。受人类认知机制启发——在推理时会回顾视觉片段,我们提出VTI-CoT,一种视觉-文本交错式思维链框架。该框架将文本推理步骤与对应视觉帧进行融合。由于现有数据集中缺乏视觉-文本交错式思维链,我们开发自动化标注流程,构建高质量多模态CoT数据。此外,长视频推理需生成越来越长的CoT token序列,严重阻碍训练收敛与效率。为此,我们采用基于OCR的压缩技术,将CoT监督信号压缩至单一画布。实验表明,VTI-CoT在同等参数规模模型中达到当前最优性能,且显著提升训练效率。

原文摘要 · Abstract (English)

Video reasoning aims to understand complex temporal events and causal relationships within videos. Recently, Chain-of-Thought (CoT) has been introduced to this field to enhance reasoning accuracy. However, existing CoT-based video reasoning methods primarily rely on text-only information for logical deduction, overlooking critical visual information during the inference process. Inspired by the human cognitive mechanism of reviewing visual segments during inference, we propose VTI-CoT, a Visual-Textual Interleaved CoT framework. VTI-CoT integrates textual reasoning steps with corresponding visual frames. Given the scarcity of visual-textual interleaved CoT in existing datasets, we develop an automated annotation pipeline to construct high-quality multimodal CoT data. Further, reasoning over long-form videos entails increasingly long CoT token sequences, which severely hinders training convergence and efficiency. To address this, we employ Optical Character Recognition (OCR)-based compression techniques to compress CoT supervision signals into a single canvas. Experimental results demonstrate that VTI-CoT achieves state-of-the-art performance among models of the same parameter scale while significantly improving training efficiency.

视频推理思维链多模态OCR压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。