arXiv:2510.13149cs.RO2025-10被引 9

提出新评估框架,测试机器人在复杂扰动下组合技能的能力

RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation

  • 设计分层评估范式,分离规划与执行以检验技能组合能力
  • 在多种扰动下,现有模型组合泛化能力明显不足
  • 适合研究机器人长期操作与技能组合的学者参考

让机器人在多样扰动下灵活调度和组合已学技能以完成新型长时序操作任务仍是核心挑战。早期端到端视觉-语言-动作(VLA)模型表现有限,难以超越训练分布。分层方法虽通过高层规划生成子目标给低层策略带来一定改进,但在复杂扰动下仍显不足,暴露出技能组合能力受限。然而,现有基准主要关注长时序任务的完成率,缺乏对组合泛化、鲁棒性及规划-执行交互的深入分析。为此,我们提出 RoboHiMan,一个用于长时序操作中组合泛化的能力评估框架。该框架引入 HiMan-Bench 基准,包含原子任务与组合任务,在多种扰动下进行评估,并配套多层级训练数据集用于分析数据规模的影响;同时提出三种评估范式(常规、解耦、耦合),用以探查技能组合的必要性并揭示分层架构中的瓶颈。实验表明,代表性模型在不同架构下均存在明显能力差距,指明了提升真实世界长时序操作模型的方向。视频与开源代码见项目主页:https://chenyt31.github.io/robo-himan.github.io/

原文摘要 · Abstract (English)

Enabling robots to flexibly schedule and compose learned skills for novel long-horizon manipulation under diverse perturbations remains a core challenge. Early explorations with end-to-end VLA models show limited success, as these models struggle to generalize beyond the training distribution. Hierarchical approaches, where high-level planners generate subgoals for low-level policies, bring certain improvements but still suffer under complex perturbations, revealing limited capability in skill composition. However, existing benchmarks primarily emphasize task completion in long-horizon settings, offering little insight into compositional generalization, robustness, and the interplay between planning and execution. To systematically investigate these gaps, we propose RoboHiMan, a hierarchical evaluation paradigm for compositional generalization in long-horizon manipulation. RoboHiMan introduces HiMan-Bench, a benchmark of atomic and compositional tasks under diverse perturbations, supported by a multi-level training dataset for analyzing progressive data scaling, and proposes three evaluation paradigms (vanilla, decoupled, coupled) that probe the necessity of skill composition and reveal bottlenecks in hierarchical architectures. Experiments highlight clear capability gaps across representative models and architectures, pointing to directions for advancing models better suited to real-world long-horizon manipulation tasks. Videos and open-source code can be found on our project website: https://chenyt31.github.io/robo-himan.github.io/.

机器人操作组合泛化分层控制评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。