arXiv:2603.12746cs.CV2026-03被引 14

评测大模型对动态物理世界的感知与推理能力,发现现有模型在运动理解上表现不一。

Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

  • 构建动态视频基准集Dyn-Bench,包含1000个视频和7000个问答对
  • 现有模型在时空推理与动态物体定位上难以同时保持高准确率
  • 引入结构化融合方法显著提升模型对动态变化的理解能力

人类生活在具有时空维度的物理4D世界中,其几何结构与语义内容随时间演变。尽管当前多模态大语言模型(MLLMs)在静态视觉理解上表现优异,但它们能否具备“动态思维”——即感知、追踪并推理复杂场景中的时空动态?为系统评估其时空推理与局部动态感知能力,我们提出了一个大规模基准测试Dyn-Bench,基于多样真实与合成视频数据构建,支持对时空理解能力的稳健、可扩展评估。通过从海量2D与4D数据源进行多阶段筛选,Dyn-Bench提供高质量动态场景集合,包含1000个视频、7000个视觉问答(VQA)对及3000个动态物体定位对。我们对通用型、空间级和区域级的MLLMs进行探查,分析其在语言与视觉层面如何表达对动态的认知,结果发现现有模型无法同时在时空推理与动态物体定位上保持强性能,常对运动与交互产生不一致解释。值得注意的是,传统提示策略(如思维链或描述性提示)改善有限,而结构化整合方法(包括掩码引导融合与时空文本认知图,ST-TCM)则显著增强模型在物理4D世界中的动态感知与时空推理能力。代码与基准已公开于https://dyn-bench.github.io/。

原文摘要 · Abstract (English)

Humans inhabit a physical 4D world where geometric structure and semantic content evolve over time, constituting a dynamic 4D reality (spatial with temporal dimension). While current Multimodal Large Language Models (MLLMs) excel in static visual understanding, can they also be adept at "thinking in dynamics", i.e., perceive, track and reason about spatio-temporal dynamics in evolving scenes? To systematically assess their spatio-temporal reasoning and localized dynamics perception capabilities, we introduce Dyn-Bench, a large-scale benchmark built from diverse real-world and synthetic video datasets, enabling robust and scalable evaluation of spatio-temporal understanding. Through multi-stage filtering from massive 2D and 4D data sources, Dyn-Bench provides a high-quality collection of dynamic scenes, comprising 1k videos, 7k visual question answering (VQA) pairs, and 3k dynamic object grounding pairs. We probe general, spatial and region-level MLLMs to express how they think in dynamics both linguistically and visually, and find that existing models cannot simultaneously maintain strong performance in both spatio-temporal reasoning and dynamic object grounding, often producing inconsistent interpretations of motion and interaction. Notably, conventional prompting strategies (e.g., chain-of-thought or caption-based hints) provide limited improvement, whereas structured integration approaches, including Mask-Guided Fusion and Spatio-Temporal Textual Cognitive Map (ST-TCM), significantly enhance MLLMs' dynamics perception and spatio-temporal reasoning in the physical 4D world. Code and benchmark are available at https://dyn-bench.github.io/.

多模态动态推理视频理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。