arXiv:2602.20687cs.AI2026-02AAAI被引 2

新基准揭示视觉语言模型在具身智能中的基础能力短板

How Foundational Skills Influence VLM-based Embodied Agents:A Native Perspective

  • 构建统一低级动作空间的原生环境,更贴近真实控制
  • 发现顶级视觉语言模型在基础操作技能上存在明显缺陷
  • 适合研究具身智能与多粒度评估的学者参考

近期视觉语言模型(VLMs)在实现类人具身智能方面展现出潜力。然而,现有针对VLM驱动的具身智能体的基准测试通常依赖高层指令或离散动作空间,属于非原生设置,与真实世界控制差异显著。此外,当前基准主要关注高层任务,缺乏对低层与高层任务的联合评估与分析。为解决这些问题,我们提出NativeEmbodied,一个基于多样化仿真场景的挑战性基准,采用统一的原生低级动作空间。该基准包含三个复杂场景下的代表性高层任务,用于评估整体表现;为进一步分析,我们拆解复杂任务所需技能,构建四类低级任务,分别对应四种基础具身技能。跨任务与技能粒度的联合评估实现了对具身智能体的细粒度测评。使用先进VLMs的实验揭示了多个基础具身技能的明显不足,进一步分析表明这些瓶颈严重制约高层任务表现。NativeEmbodied揭示了当前VLM驱动具身智能体的关键挑战,并为未来研究提供指导。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have shown promise for human-level embodied intelligence. However, existing benchmarks for VLM-driven embodied agents often rely on high-level commands or discretized action spaces, which are non-native settings that differ markedly from real-world control. In addition, current benchmarks focus primarily on high-level tasks and lack joint evaluation and analysis at both low and high levels. To address these limitations, we present NativeEmbodied, a challenging benchmark for VLM-driven embodied agents that uses a unified, native low-level action space. Built on diverse simulated scenes, NativeEmbodied includes three representative high-level tasks in complex scenarios to evaluate overall performance. For more detailed analysis, we further decouple the skills required by complex tasks and construct four types of low-level tasks, each targeting a fundamental embodied skill. This joint evaluation across task and skill granularities enables fine-grained assessment of embodied agents. Experiments with state-of-the-art VLMs reveal clear deficiencies in several fundamental embodied skills, and further analysis shows that these bottlenecks significantly limit performance on high-level tasks. NativeEmbodied highlights key challenges for current VLM-driven embodied agents and provides insights to guide future research.

具身智能视觉语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。