arXiv:2510.07550cs.CVcs.AI2025-10被引 13

让视频语言模型更懂物理常识,识别生成视频中的荒谬动作

TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility

  • 用轨迹感知注意力机制增强模型对运动规律的捕捉能力
  • 在300个视频的基准测试中,新方法使判断准确率提升21%
  • 适合研究多模态模型物理理解能力的研究者和开发者

尽管现代视频生成模型具有出色的视觉保真度,但常产生违背直觉物理规律的序列,如物体悬浮、瞬移或非因果变形。人类能轻松察觉此类异常,但尚无可靠方法量化评估视频的物理真实性。本文探索视频语言模型(VLMs)作为物理合理性裁判的潜力,发现现有VLMs在时空与因果推理方面存在根本缺陷。为此,提出TRAVL微调方案,结合平衡数据集与轨迹感知注意力模块,提升模型对运动编码与差异的敏感性。为更严格评估物理推理能力,构建ImplausiBench基准,包含300个视频(150个真实,150个生成),消除语言偏见并隔离视觉-时间理解。性能通过人工金标准与更严格的LLM作为裁判指标报告。TRAVL与ImplausiBench共同构成统一框架,用于探测和改进多模态模型的物理合理性,揭示视觉-时间理解中一个关键且未充分探索的挑战。

原文摘要 · Abstract (English)

Despite impressive visual fidelity, modern video generative models frequently produce sequences that violate intuitive physical laws, such as objects floating, teleporting, or morphing in ways that defy causality. While humans can easily detect such implausibilities, there remains no robust method for quantitatively assessing physical realism in video. In this work, we explore whether Video-Language Models (VLMs) can be trained to serve as reliable judges of physical plausibility. We find that existing VLMs struggle to identify physics violations, exposing fundamental limitations in their temporal and causal reasoning. To address this, we introduce TRAVL, a fine-tuning recipe that combines a balanced training dataset with a trajectory-aware attention module to improve motion encoding and discrimination in VLMs. To evaluate physical reasoning more rigorously, we propose ImplausiBench, a benchmark of 300 videos (150 real, 150 generated) that removes linguistic biases and isolates visual-temporal understanding. Performance is reported both with gold-standard human judgments and stricter LLM-as-judge metrics. Together, TRAVL and ImplausiBench offer a unified framework for probing and improving physical plausibility in multimodal models, shedding light on a challenging and underexplored aspect of visual-temporal understanding.

视频生成物理推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。