arXiv:2605.04515cs.CV2026-05

让视频大模型学会基于视觉事实推理物理规律,解决错觉与反直觉问题。

From Priors to Perception: Grounding Video-LLMs in Physical Reality

论文配图:From Priors to Perception: Grounding Video-LLMs in Physical Reality
图 1 · 摘自论文原文
  • 用物理定律生成对抗性视频数据,分离视觉幻觉与逻辑错误。
  • 通过视觉锚定推理链,显著提升模型在真实物理场景中的判断力。
  • 无需修改架构,仅用轻量微调即可突破现有模型的语义先验限制。

尽管视频大语言模型在通用理解上表现优异,但在细粒度物理推理方面存在系统性缺陷。现有方法不仅泛化能力有限,且混淆生成幻象与真实的物理谬误。研究发现,模型不仅在违反物理规律的异常场景中失败,也在视觉事实与统计预期相悖的反直觉情境中出错。为此,提出统一归因理论:此类双重失败并非感知缺陷,而是由语义先验主导所致——模型的推理被内部叙事脚本深度干扰。为此构建首个基于物理定律的高保真对抗性视频数据集PACC,彻底解耦视觉伪影与逻辑错误。同时设计视觉锚定推理链(VARC),强制模型在逻辑判断前显式依赖低层视觉事实。实验表明,无需侵入式架构修改,仅通过PACC进行标准LoRA微调,即可有效消除先验干扰,在主流(SOTA)模型上实现物理推理能力的显著跃升。

原文摘要 · Abstract (English)

While Video Large Language Models (Video-LLMs) excel in general understanding, they exhibit systematic deficits in fine-grained physical reasoning. Existing interventions not only suffer from limited generalization but fundamentally conflate generative artifacts with genuine physical fallacies. Furthermore, we find that models fail systematically not only in anti-physics anomalies but also in counter-intuitive scenarios where visual facts contradict statistical expectations. Accordingly, we propose the Unified Attribution Theory: this dual failure stems not from perception deficiency, but from Semantic Prior Dominance -- the reasoning mechanism is deeply hijacked by internal narrative scripts. To address this, we construct the Programmatic Adversarial Curriculum (PACC), the first high-fidelity adversarial video dataset synthesized based on physical laws, thoroughly decoupling visual artifacts from logical errors. Concurrently, we design the Visual-Anchored Reasoning Chain (VARC) to force models to explicitly ground their judgments in low-level visual facts prior to logical adjudication. Experiments demonstrate that without invasive architectural modifications, standard LoRA fine-tuning with the PACC curriculum effectively neutralizes prior interference in state-of-the-art (SOTA) models, yielding a substantial leap in physical reasoning capabilities.

视频理解物理推理因果建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。