视频转可执行世界,零训练实现物理推理突破
PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning

- 从视频构建可复用的动态场景,无需训练
- 在CLEVRER上比直接推理高38.23分,反事实任务超基线19.25分
- 适合需要精准物理推理的科研与工业应用
可靠地从视频中进行物理推理,需理解物体运动、交互及对干预的响应。现有视觉语言模型(VLM)常难以解析此类动态并可靠推断未来或反事实结果。我们提出PhysMind,一种零训练的智能体框架,为每段视频构建一个可重用、问题无关的可执行世界。PhysMind通过对象分割、网格重建和6D姿态跟踪恢复时间一致的动态场景,再拟合连续时间动力学与潜在物理参数,无需步进式模拟器。给定问题后,它可检查、延续或修改世界,并从生成轨迹与交互中得出答案。相比同源VLM的直接链式思维(CoT)推理,PhysMind在CLEVRER上提升38.23分,在Physion++上提升8.08分;在反事实问题上,超过最强基线GPT-5.5达19.25分。
原文摘要 · Abstract (English)
Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language models (VLMs) often struggle to interpret these dynamics and reason reliably about future and counterfactual outcomes. We introduce PhysMind, a training-free agentic framework that constructs one reusable, question-agnostic executable world per video. PhysMind recovers a temporally consistent dynamic scene through object segmentation, mesh reconstruction, and 6D pose tracking, then fits analytic continuous-time dynamics and latent physical parameters without unrolling a time-stepped simulator. Given a question, it inspects, continues, or edits the world and answers from the resulting trajectories and interactions. Relative to direct chain-of-thought (CoT) reasoning with the same VLM, PhysMind improves accuracy by 38.23 points on CLEVRER and 8.08 points on Physion++. On counterfactual questions, it exceeds the strongest evaluated VLM baseline, GPT-5.5, by 19.25 points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。