测试视觉模型对车辆运动的物理理解能力,发现其视觉与逻辑脱节。
EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving

- 用确定性映射将运动轨迹转为离散概念,分离视觉与物理逻辑
- 20多个模型均表现不佳,甚至不如传统几何方法
- 语言模态主导运动推理,视觉贡献几乎为零
尽管视觉语言模型在自动驾驶中提升了高层推理能力,但其对自身推理所依赖的本体运动物理基础的理解仍不清晰。我们提出EgoDyn-Bench,一个用于评估视觉中心型基础模型语义本体运动理解能力的诊断基准。通过确定性归因器将连续车辆运动学映射为离散运动概念,我们将模型内部物理逻辑与视觉感知解耦。大规模实证审计涵盖20+模型,包括闭源多模态大模型、多尺度开源视觉语言模型及专用视觉语言架构,揭示显著的感知瓶颈:尽管模型具备逻辑上的物理概念,却无法准确与视觉观测对齐,频繁劣于经典非学习几何基线。该失败在不同模型规模和领域训练下持续存在,表明当前架构在视觉感知与物理推理耦合上存在结构性缺陷。我们证明,提供显式轨迹编码可显著恢复所有模型的物理一致性,揭示视觉与语言间功能解耦:本体运动逻辑几乎完全来自语言模态,而视觉观察提供的时序信号微乎其微。这一结构发现为可解释的具身智能提供了标准化诊断框架与可行路径。
原文摘要 · Abstract (English)
While Vision-Language Models (VLMs) have advanced high-level reasoning in autonomous driving, their ability to ground this reasoning in the underlying physics of ego-motion remains poorly understood. We introduce EgoDyn-Bench [Project page: (https://tum-avs.github.io/EgoDyn-Bench-Website/), Code: (https://github.com/TUM-AVS/EgoDyn-Bench), Dataset: (https://huggingface.co/datasets/fnc1901/EgoDyn-Bench)], a diagnostic benchmark for evaluating the semantic ego-motion understanding of vision-centric foundation models. By mapping continuous vehicle kinematics to discrete motion concepts via a deterministic oracle, we decouple a model's internal physical logic from its visual perception. Our large-scale empirical audit spanning 20$+$ models, including closed-source MLLMs, open-source VLMs across multiple scales, and specialized VLAs, identifies a significant Perception Bottleneck: while models exhibit logical physical concepts, they consistently fail to accurately align them with visual observations, frequently underperforming classical non-learned geometric baselines. This failure persists across model scales and domain-specific training, indicating a structural deficit in how current architectures couple visual perception with physical reasoning. We demonstrate that providing explicit trajectory encodings substantially restores physical consistency across all evaluated models, revealing a functional disentanglement between vision and language: ego-motion logic is derived almost exclusively from the language modality, while visual observations contribute negligible temporal signal. This structural finding provides a standardized diagnostic framework and a practical pathway toward physically aligned embodied AI. Ego-motion - Physical Reasoning - Foundation Models
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。