arXiv:2605.20576cs.CV2026-05中稿 · CVPR被引 1

用语言描述物理运动,让模型从视频中推断刚体动态。

$Δ$ynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos

论文配图:$Δ$ynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos
图 1 · 摘自论文原文
  • 用自然语言结构化表达刚体状态,替代直接预测参数。
  • 在CLEVRER上分割IoU达0.30,比顶尖视觉语言模型高7倍。
  • 适合需要跨场景物理推理的仿真与感知研究者。

从单目视频中推断刚体物理状态和属性是实现基于物理的感知与模拟的关键步骤。现有方法依赖特定物理系统、物体类型和相机姿态,难以泛化到复杂真实场景。我们提出ΔYNAMICS,一个以语言为统一表征的视觉-语言框架,将刚体动力学用结构化文本形式生成,用于物理模拟。通过引入自然语言运动推理并利用光流作为语义无关输入,提升模型泛化能力。在CLEVRER数据集上,ΔYNAMICS分割交并比(IoU)达到0.30,相比InternVL3-8B、Qwen2.5-VL-7B和Claude-4-Sonnet等领先VLMs提升7倍。此外,测试时采样和进化搜索分别使性能再提升27%和120%。最后,我们在包含235个真实世界刚体视频的新数据集上验证了良好迁移性,展示了语言驱动物理推断在连接感知与模拟方面的潜力。

原文摘要 · Abstract (English)

Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses, making them unable to generalize to complex real-world settings. We introduce $Δ$YNAMICS, a vision-language framework that uses language as a unified representation of rigid-body dynamics. Instead of directly predicting parameters, $Δ$YNAMICS generates scene configurations in a structured text format for physics simulation. We enhance the model's generalization by integrating natural language motion reasoning and leveraging optical flow as a semantic-agnostic input. On the CLEVRER dataset, $Δ$YNAMICS achieves a segmentation IoU of 0.30, a 7x improvement over leading VLMs (InternVL3-8B, Qwen2.5-VL-7B and Claude-4-Sonnet). Additionally, test-time sampling and evolutionary search further boost performance by 27% and 120% in segmentation IoU, respectively. Finally, we demonstrate strong transfer to a new dataset of 235 real-world rigid-body videos, highlighting the potential of language-driven physics inference for bridging perception and simulation.

物理推断视觉语言刚体动力学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。