arXiv:2512.18028cs.ROcs.CV2025-12被引 3

评测视觉语言模型在真实机器人平台上的综合推理能力。

Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation

  • 设计闭环测试,用三类机器人平台评估模型跨平台推理能力。
  • 1.1K问答+58导航任务,检验语义、空间、时间、物理四维推理。
  • 发现空间与时间推理是当前模型最大短板,跨模态对齐更关键。

视觉语言导航要求智能体在具身约束下进行推理与行动。尽管视觉语言模型(VLMs)表现出强泛化能力,但现有基准难以揭示具身性——即物理平台选择、传感器配置与模态对齐——如何影响感知、推理与控制。我们提出Embodied4C,一个闭合回路基准,作为具身推理的图灵测试。该基准通过约1.1K个一次性推理问题和58个目标导向导航任务,评估十种状态最优的VLMs在三种异构具身平台(自动驾驶车辆、航拍无人机、机械臂)上的表现。任务联合评估语义、空间、时间与物理四维基础能力。每种平台均配置动态传感器与环境变化,以探测超越平台特异性适应的泛化能力。为防止具身过拟合,引入跨域远距离查询,考察抽象与跨上下文推理。对十种先进VLMs与四种具身控制基线的全面评估表明:跨模态对齐与指令微调比模型规模更重要,而空间与时间推理仍是实现可靠具身能力的主要瓶颈。

原文摘要 · Abstract (English)

Vision-language navigation requires agents to reason and act under constraints of embodiment. While vision-language models (VLMs) demonstrate strong generalization, current benchmarks provide limited understanding of how embodiment -- i.e., the choice of physical platform, sensor configuration, and modality alignment -- influences perception, reasoning, and control. We introduce Embodied4C, a closed-loop benchmark designed as a Turing test for embodied reasoning. The benchmark evaluates the core embodied capabilities of VLMs across three heterogeneous embodiments -- autonomous vehicles, aerial drones, and robotic manipulators -- through approximately 1.1K one-shot reasoning questions and 58 goal-directed navigation tasks. These tasks jointly assess four foundational dimensions: semantic, spatial, temporal, and physical reasoning. Each embodiment presents dynamic sensor configurations and environment variations to probe generalization beyond platform-specific adaptation. To prevent embodiment overfitting, Embodied4C integrates domain-far queries targeting abstract and cross-context reasoning. Comprehensive evaluation across ten state-of-the-art VLMs and four embodied control baselines shows that cross-modal alignment and instruction tuning matter more than scale, while spatial and temporal reasoning remains the primary bottleneck for reliable embodied competence.

具身智能视觉语言多模态导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。