arXiv:2604.03302cs.CVcs.AI2026-04

提出场景动态场方法,提升多模态大模型对物理动态的理解能力

Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models

论文配图:Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
图 1 · 摘自论文原文
  • 引入物理模拟器,在多任务微调中构建场景动态场
  • 在流体任务上性能提升20.7%,且泛化到未见物理场景
  • 适合关注物理推理与多模态理解的科研人员

尽管多模态大语言模型(MLLMs)在图像和视频理解方面表现优异,但其对物理世界的理解仍面临挑战。本文聚焦于直观物理理解这一基础环节,发现现有模型在连续物体动态理解上存在显著不足。为此,我们设计了两个基准任务:下一帧选择(NFS)和时间一致性验证(TCV),实验表明即使是顶尖的MLLMs在这些任务上表现依然不佳。为解决该问题,我们提出场景动态场(SDF),一种结合物理模拟器的多任务微调框架。SDF显著提升模型性能,在流体任务上最高实现20.7%的提升,并展现出对未见物理领域的强泛化能力。本工作不仅揭示了当前MLLMs在物理理解上的关键短板,也提供了一种高效低成本的改进路径。代码与数据已开源。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-level physics reasoning. In this work, we investigate the first step of physical reasoning, i.e., intuitive physics understanding, revealing substantial limitations in understanding the dynamics of continuum objects. To isolate and evaluate this specific capability, we introduce two fundamental benchmark tasks: Next Frame Selection (NFS) and Temporal Coherence Verification (TCV). Our experiments demonstrate that even state-of-the-art MLLMs perform poorly on these foundational tasks. To address this limitation, we propose Scene Dynamic Field (SDF), a concise approach that leverages physics simulators within a multi-task fine-tuning framework. SDF substantially improves performance, achieving up to 20.7% gains on fluid tasks while showing strong generalization to unseen physical domains. This work not only highlights a critical gap in current MLLMs but also presents a promising cost-efficient approach for developing more physically grounded MLLMs. Our code and data are available at https://github.com/andylinx/Scene-Dynamic-Field.

多模态物理推理动态理解模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。