提出新基准,测试多模态模型是否真懂物理世界。
ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?

- 用历史视频和文本推断未来物理状态,强制跨模态推理
- 1万+长视频数据集,500万词标注,暴露模型依赖语言捷径
- 发现主流模型仍缺乏真实物理推理能力,适合研究可信AI的学者
近期多模态大模型在开放世界推理中表现突出,但其是否真正融合多模态信息进行物理化推理仍存疑。为消除语言模态偏差与捷径依赖,本文提出新基准 ChronoPhyBench,将下一状态预测与视觉问答结合,基于历史视频上下文和文本描述,要求模型通过单图选择或多帧时序排序来推断后续物理状态。同时构建大规模数据集,包含超10,000条长视频与精标注文本,总计500万词。实验表明,现有开源模型在物理基推理方面仍处于初级阶段,远未达到真实跨模态理解水平。本工作旨在系统评估多模态模型推理能力,量化幻觉率,推动物理人工智能发展,为通用人工智能提供透明可靠的评估框架。
原文摘要 · Abstract (English)
Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in open-world reasoning and understanding. However, a critical ambiguity persists: it remains unclear whether these models genuinely synthesize cross-modal information to construct physically grounded reasoning chains, or if they merely exploit strong language priors to mask single-modality reliance, thereby hallucinating advanced multimodal capabilities. Motivated by this, and to rigorously mitigate language modality bias and shortcuts, we propose a novel multimodal Chrono}logical Physical Dynamics Reasoning Benchmark ChronoPhyBench, which unifies next state prediction with Visual Question Answering (VQA) paradigms by conditioning on historical video context and textual captions to enforce models to deduce subsequent physical states through both single image selection and the inherently more complex task of multiple frame chronological sorting. Concurrently, we construct a large-scale multimodal reasoning dataset curated using the ChronoPhyBench criteria, comprising over 10,000 long-form videos paired with meticulously annotated captions, totaling 5M tokens. Our experimental evaluations reveal a stark contrast to conclusions drawn by previous benchmarks. The capacity of current open-source models to perform physically grounded multimodal reasoning remains in its infancy. Ultimately, this work seeks to systematically stress-test the reasoning capabilities of multimodal models, quantify hallucination rates, and advance the development of Physical AI, thereby providing the community with a robust and transparent evaluation framework toward Artificial General Intelligence (AGI).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。