arXiv:2507.09876cs.CVcs.AI2025-07中稿 · ACM MM 2025被引 26

让大模型像人一样边看视频边思考,提升视频理解能力

ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models

  • 视频与文本交替推理,模拟人类思考过程
  • 在新构建的基准上性能显著优于纯文本推理
  • 适合研究多模态大模型与智能体认知推理的学者

视频理解在连接低层视觉信号与高层认知推理方面至关重要,是自动驾驶、具身智能及通用人工智能发展的基础。尽管大型语言模型(LLMs)尤其是采用思维链(CoT)技术的模型在视频推理方面取得进展,但现有方法主要依赖文本信息,忽视了实际推理中对视觉内容的反复观察。受人类自然在推理时回看视觉内容的启发,本文提出一种新型视频推理范式:视频-文本交错思维链(ViTCoT),使推理更符合认知规律。为此,我们构建了视频-文本交错基准(ViTIB),通过多模态大模型筛选关键视频片段并人工验证。大量实验表明,ViTCoT显著优于传统纯文本思维链,且能更有效地激活多模态大模型中的神经元。

原文摘要 · Abstract (English)

Video understanding plays a vital role in bridging low-level visual signals with high-level cognitive reasoning, and is fundamental to applications such as autonomous driving, embodied AI, and the broader pursuit of AGI. The rapid development of large language models (LLMs), particularly those utilizing Chain-of-Thought (CoT) technology, has significantly advanced video reasoning capabilities. However, current approaches primarily depend on textual information for reasoning, overlooking the visual modality in the actual video reasoning process. In contrast, humans naturally re-examine visual content while reasoning. Motivated by this, we introduce a novel video reasoning paradigm: Video-Text Interleaved CoT (ViTCoT), which facilitates more intuitive and cognitively aligned reasoning. To the end, first, we construct the Video-Text Interleaved Benchmark (ViTIB), which is created using MLLMs for key-video selection and manually verified. Furthermore, we extensively explore the potential of the ViTCoT paradigm in the video understanding field. Extensive experiments demonstrate that ViTCoT significantly enhances performance compared to the traditional text-only CoT paradigm and effectively activates more neuron values in MLLMs.

视频理解多模态思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。