arXiv:2510.08759cs.CVcs.RO2025-10中稿 · ICML被引 4

通过分解技能评估,揭示多模态大模型在具身任务中的瓶颈

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

  • 将具身任务拆解为14项基础技能,实现细粒度能力评估
  • 发现感知能力是推理失败的主要瓶颈,时空建模不稳定
  • 提出BEAR-Agent增强模型视觉空间推理,实测提升17.5%

理解具身多模态大语言模型(MLLMs)的能力瓶颈对提升具身智能体至关重要。然而,现有具身基准主要关注任务级评估,难以提供模型失败的深层原因。为此,我们引入BEAR基准,将具身任务分解为14项原子技能,进行细粒度技能级评估。BEAR包含4,469个交错的图像-视频-文本样本,覆盖6类共14项技能,从低层感知到高层规划。我们在层次化技能诊断框架下评估了20个MLLMs,发现:(1) 感知能力是推理失败的主要瓶颈;(2) 当前模型存在不稳定的时空建模问题,此前基准未暴露此缺陷。基于此,我们提出BEAR-Agent,一种融合视觉与空间推理工具的多模态对话代理。BEAR-Agent在具身技能上显著提升性能,在GPT-5上相对基线提升17.5%,并在仿真与真实机器人实验中超越强基线。

原文摘要 · Abstract (English)

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, existing embodied benchmarks mainly focus on task-level evaluation and fail to provide actionable insights into the underlying causes of model failures. To address this limitation, we introduce BEAR, a benchmark that decomposes embodied tasks into 14 atomic skills for fine-grained skill-level evaluation. BEAR comprises 4,469 interleaved image-video-text samples spanning 14 skills across 6 categories, ranging from low-level perception to high-level planning. We evaluate 20 MLLMs on BEAR under a hierarchical skill-level diagnosis framework and uncover two key findings: (1) perceptual capabilities are major bottlenecks behind reasoning failures, and (2) current models suffer from unstable spatiotemporal modeling that remains largely unexposed in prior benchmarks. Motivated by these findings, we further propose BEAR-Agent, a multimodal conversational agent that augments MLLMs with visual and spatial reasoning tools. BEAR-Agent substantially improves performance across embodied skills, achieving a relative improvement of 17.5% on GPT-5 over the base model on BEAR, while also outperforming strong baselines in both simulation and real-world robotic experiments. Project page: https://bear-official66.github.io/

具身智能多模态模型技能评估机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。