arXiv:2607.10190cs.LGcs.AI2026-07

用结构化物理记忆提升视频模型对物理合理性判断的能力

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning

论文配图:PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning
图 1 · 摘自论文原文
  • 构建分层物理记忆库,包含场景、事件关系和物理规则三类知识
  • 在不微调模型的前提下,使多个主流视频模型在3个基准上表现显著提升
  • 适合需要增强物理推理能力的视觉语言模型应用

视频-语言模型(VLM)在视频理解与视觉问答中表现优异,但在物理合理性推理方面仍不可靠,尤其在复杂物理推理基准上暴露出现有模型在物理常识推理上的持续短板。为解决该问题,我们提出 PhysMRV——一种无需训练的物理记忆检索与验证框架。不同于仅检索语义相似视频的增强方法,PhysMRV将训练视频转化为分层记忆库,包含三类互补知识:场景描述(捕捉视觉上下文)、物理事件图(建模物体交互与因果结构)、物理规则摘要(提炼可复用的物理原理与线索)。推理时,PhysMRV检索相关物理记忆,并利用其结构化物理证据引导冻结的VLM进行物理合理性验证,无需微调或参数更新。我们在ImplausiBench、IntPhys2和GRASP Level 2三个挑战性物理推理基准上,对多种主流VLM进行了评估。结果表明,相比直接提示,PhysMRV在各类VLM与基准上均实现稳定提升,证明结构化物理记忆是无需额外训练即可有效增强物理合理性推理的可扩展路径。

原文摘要 · Abstract (English)

Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plausibility, where understanding object interactions, causal dynamics, and fundamental physical principles is essential. This limitation is particularly evident on challenging physical reasoning benchmarks, revealing a persistent gap in physical commonsense reasoning. To address this challenge, we propose PhysMRV, a training-free physical memory and verification framework for physical plausibility reasoning. Unlike retrieval-augmented VLMs that retrieve semantically similar videos as additional context, PhysMRV transforms training videos into a Hierarchical Memory Bank of structured physical knowledge comprising three complementary levels: scene descriptions capturing visual context, physical-event graphs modeling object interactions and causal structure, and physics-rule summaries distilling reusable physical principles and cues. During inference, PhysMRV retrieves physically relevant memories and leverages their structured physical evidence to guide a frozen VLM in verifying physical plausibility, requiring neither fine-tuning nor parameter updates. We evaluate PhysMRV on three challenging physical reasoning benchmarks, ImplausiBench, IntPhys2, and GRASP Level 2, across multiple state-of-the-art VLMs. Experimental results demonstrate consistent improvements over direct prompting across diverse VLMs and evaluation benchmarks, showing that structured physical memories provide an effective and scalable means of enhancing physical plausibility reasoning without additional training.

物理推理记忆机制视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。