arXiv:2606.16826cs.ROcs.AI2026-06

评测机器人操作策略的原子技能与组合泛化能力,发现当前模型仍难精细执行任务。

ATOM-Bench: A Real-World Benchmark for Atomic Skills and Compositional Generalization in Manipulation Policies

论文配图:ATOM-Bench: A Real-World Benchmark for Atomic Skills and Compositional Generalization in Manipulation Policies
图 1 · 摘自论文原文
  • 将操作任务拆解为动作原子和指令原子,构建真实世界测试基准
  • 2700次物理实验显示:模型在计数和逻辑过滤上表现差,组合任务失败率高
  • 提出新评估指标,区分执行失误与组合能力不足,适合研究泛化瓶颈

通用操作策略被视为机器人控制的基础模型,但其实用泛化能力难以评估。一个策略可能在演示任务中成功,却无法完成精细的原子技能或在新任务结构中复用已学技能。我们提出 extbf{ATOM-Bench},一个面向真实世界操作策略的原子技能与组合泛化评测基准。该基准将桌面操作分解为运动原子和指令原子,包含30个原子任务和24个保留的组合任务,覆盖单臂与双臂机器人路径。我们收集了3,000条人类示范用于原子微调,并公开示范数据与评估回放数据,支持可复现的真实世界评估。策略在原子任务上微调后,在原子技能获取与保留组合任务上进行评估。我们引入原子得分(AS)和组合失败率(CFS),以区分由弱原子技能导致的失败与组合复用受限造成的失败。通过对五种代表性操作策略进行2,700次物理回放,我们发现当前策略虽能掌握简单指令对齐技能,但在精细动作、计数和逻辑筛选上仍存在困难。更重要的是,强原子性能并不可靠地转化为保留组合任务的性能。ATOM-Bench 提供了一个诊断性测试平台,用于判断失败是源于执行薄弱、指令理解差,还是组合复用能力有限。

原文摘要 · Abstract (English)

Generalist manipulation policies are increasingly presented as foundation models for robotic control, but their real-world generalization remains difficult to diagnose. A policy may succeed on demonstrated tasks while still failing to execute fine-grained atomic skills or recombine learned skills in new task structures. We introduce \textbf{ATOM-Bench}, a real-world benchmark for evaluating both atomic skills and compositional generalization in manipulation policies. ATOM-Bench factorizes tabletop manipulation into motor atoms and instruction atoms, and contains 30 atomic tasks and 24 held-out compositional tasks across paired single-arm and dual-arm robot tracks. We collect 3,000 human demonstrations for atomic fine-tuning and release both the demonstration data and evaluation rollout data to support reproducible real-world evaluation. Policies are fine-tuned on atomic tasks and evaluated on both atomic skill acquisition and held-out compositional tasks. We further introduce Atomic Score (AS) and Compositional Failure Share (CFS) to distinguish failures caused by weak atomic skills from failures caused by limited compositional reuse. Through 2,700 physical rollouts on five representative manipulation policies, we find that current policies can acquire simple instruction-grounding skills, but still struggle with fine-grained motor atoms, counting, and logical filtering. More importantly, strong atomic performance does not reliably transfer to held-out compositional tasks. ATOM-Bench provides a diagnostic testbed for studying whether failures arise from weak motor execution, poor instruction grounding, or limited compositional reuse.

机器人操作组合泛化基准评测原子技能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。