arXiv:2606.18239cs.RO2026-06被引 3

EBench通过多维度诊断,揭示通用移动操作模型的真实能力差异。

EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies

论文配图:EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies
图 1 · 摘自论文原文
  • 设计26项任务,从5个能力维度和4个泛化维度评估模型表现
  • 发现高成功率模型在不同任务上能力分布差异显著,如XVLA擅长原子技能
  • 提供可量化模型短板的诊断信号,适合算法迭代与模型优化

我们提出EBench,一个用于诊断通用移动操作策略的仿真基准,超越单一成功率指标。EBench包含26个多样且具有挑战性的操作任务,按5个能力维度和4个泛化维度进行标注。我们评估了当前最先进的通用操作模型:$π_0$、$π_{0.5}$、XVLA和InternVLA-A1。结果显示,尽管各模型整体成功率相近,但能力分布迥异:$π_{0.5}$在测试成功率和训练-测试保持性上最优;InternVLA-A1在移动操作中占优,但在灵巧操作任务上表现崩溃;XVLA在一组独立的原子技能上优于其他策略。此外,EBench从4个代表性视角分析泛化能力,识别出不同分布偏移因素的影响。结果揭示了单一总分背后模型的真实优劣势。我们希望该基准能为通用操作模型的迭代提供丰富的诊断信号。

原文摘要 · Abstract (English)

We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single success-rate scalar. EBench comprises 26 diverse and challenging manipulation tasks annotated along 5 capability dimensions and 4 generalization dimensions. We evaluate state-of-the-art generalist manipulation models including $π_0$, $π_{0.5}$, XVLA, and InternVLA-A1, and reveal that models with near success rates exhibit strikingly different capability profiles: $π_{0.5}$ achieves the highest test success rate and the best train--test retention, whereas InternVLA-A1 dominates mobile manipulation but collapses on dexterous tasks, and XVLA exhibits strengths on a disjoint set of atomic skills compared to other policies. Beyond capability profiling, EBench analyzes the generalization ability from 4 representative perspectives, identifying the impact of different distribution shift factors. The results reveal strengths and weaknesses of models behind an overall score. We hope this benchmark offers a broad set of diagnostic signals to guide iteration on generalist manipulation models.

机器人模型诊断移动操作基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。