arXiv:2608.02150cs.CVcs.AI2026-08

构建细粒度物理规律理解数据集,提升视频大模型对物理常识的判断能力。

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

论文配图:PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs
图 1 · 摘自论文原文
  • 分粗粒度与细粒度两级设计,评估模型对物理现象是否合规及原因理解。
  • 在细粒度任务上,模型准确率提升显著,但对隐藏因果因素仍难识别。
  • 适合研究视频大模型物理推理能力、提升其可解释性与可信度的研究者。

具身智能与世界模型需要视频理解系统超越物体和动作识别,深入理解物理规律。然而,现有视频语言模型在判断事件是否符合特定物理定律方面仍不可靠。现有基准主要评估生成视频的物理质量,难以系统评估和改进视频大模型(VideoLLMs)的物理理解能力。为此,我们提出PhyCheck,一个分两级粒度组织的视频问答数据集:粗粒度子集要求模型判断视频现象是否符合或违反物理定律;细粒度子集进一步检验模型能否捕捉导致违反或符合的具体物理细节。该数据集还包含诊断子集,提供外部因果上下文以揭示影响物理合理性的隐藏因素,评估模型是否能据此调整判断。在Fine-tune Qwen2.5-VL上的实验表明,使用该数据训练可显著提升模型对物理一致性理解,但在诊断子集上仍显示当前模型难以整合额外因果条件。这揭示了表面不一致识别与深层物理机制理解之间的差距,为评估和提升VideoLLMs的物理理解能力提供了基础。

原文摘要 · Abstract (English)

Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of physical regularities. However, despite their strong performance on general video understanding tasks, current video-language models still struggle to reliably determine whether an observed event conforms to specific physical laws. Existing benchmarks primarily assess the physical quality of generated videos, providing limited support for systematically evaluating and improving the physical-law understanding of Video Large Language Models (VideoLLMs). To address this gap, we introduce PhyCheck, a video question answering dataset organized at two complementary levels of granularity. The coarse-grained subset asks models to determine whether the phenomenon shown in a video conforms to or violates physical laws, while the fine-grained subset further examines whether models can capture physical details responsible for the violation or compliance. We use these subsets as structured supervision to improve physical understanding. In addition, the dataset contains a diagnostic subset with external causal context that reveal hidden factors affecting physical plausibility, assessing whether models can recalibrate their judgments accordingly. Experiments with Fine-tune Qwen2.5-VL show that training with the proposed data substantially improves the understanding of physical-consistency, while evaluations in the diagnostic subset reveal that current models still have difficulty incorporating additional causal conditions into their decisions. These findings highlight the gap between recognizing surface-level inconsistencies and understanding underlying physical mechanisms, and provide a foundation for evaluating and improving physical understanding in Video-LLMs.

视频理解物理推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。