用视频生成模型做多模态推理,让大模型像看视频一样思考。
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- 用视频作为统一媒介,融合文本与视觉信息进行推理。
- 在视觉任务上超越GPT-5 10%,文本任务准确率达92%(MATH)。
- 适合研究多模态推理、视频生成与通用智能的学者参考。
“以文思辨”和“以图思辨”显著提升了大语言模型(LLMs)与视觉语言模型(VLMs)的推理能力,但存在固有局限:图像仅捕捉静态瞬间,难以表达动态过程;文本与视觉分属不同模态,阻碍统一理解与生成。为此,我们提出“以视频思辨”新范式,利用Sora-2等视频生成模型,将视频帧作为统一多模态推理媒介。为支持该探索,我们构建了视频思辨基准测试集VideoThinkBench,涵盖视觉主导任务(如眼力谜题)和文本主导任务(如GSM8K、MMMU)。评估表明,Sora-2具备强大推理能力:在视觉任务上接近当前最优VLM,眼力谜题表现优于GPT-5 10%;在文本任务上,数学题准确率达92%,MMMU达69.2%。我们进一步分析其能力来源,发现自一致性与上下文学习可提升性能。结果表明,视频生成模型或可成为统一多模态理解与生成的核心载体,‘以视频思辨’或为未来多模态推理的关键范式。
原文摘要 · Abstract (English)
The "Thinking with Text" and "Thinking with Images" paradigms significantly improve the reasoning abilities of large language models (LLMs) and Vision-Language Models (VLMs). However, these paradigms have inherent limitations. (1) Images capture only single moments and fail to represent dynamic processes or continuous changes, and (2) The separation of text and vision as distinct modalities, which hinders unified multimodal understanding and generation. Therefore, we propose "Thinking with Video", a new paradigm that leverages video generation models such as Sora-2 to use video frames as a unified medium for multimodal reasoning. To support this exploration, we developed the Video Thinking Benchmark (VideoThinkBench), which covers both vision-centric tasks (e.g., Eyeballing Puzzles) and text-centric tasks (e.g., GSM8K and MMMU). Our evaluation on VideoThinkBench establishes Sora-2 as a capable reasoner. On vision-centric tasks, Sora-2 is comparable to state-of-the-art (SOTA) VLMs, and even surpasses GPT-5 by 10% on eyeballing puzzles. On text-centric tasks, Sora-2 achieves 92% accuracy on MATH, and 69.2% accuracy on MMMU. Furthermore, we systematically analyze the source of these abilities. We also find that self-consistency and in-context learning can improve Sora-2's performance. In summary, our findings show that the video generation model is the potential unified multimodal understanding and generation model, positioning "Thinking with Video" as a potential unified multimodal reasoning paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。