arXiv:2607.09234cs.RO2026-07

无需标注技能,直接从混合动作数据中学习并协调行为,实现复杂重排任务。

Implicit-Behavior Coordination from Unlabeled Sub-Task Demonstrations for Rearrangement Tasks

论文配图:Implicit-Behavior Coordination from Unlabeled Sub-Task Demonstrations for Rearrangement Tasks
图 1 · 摘自论文原文
  • 从无标签子任务演示中隐式学习行为并用价值引导选择动作
  • 在复杂任务上超越专用模仿基线,接近最优规划器性能
  • 适合大规模行为库和长序列任务,无需预定义技能边界

长时序机器人重排任务通常被当作技能序列问题处理,需预先定义技能、标签或边界,以及特定任务的切换逻辑。尽管有效,但随着行为数量和任务长度增加,这种显式技能抽象难以扩展。本文提出从无标签子任务演示中进行隐式行为协调,直接从混合行为数据中学习类技能行为,并通过价值引导的动作选择进行协调。在Habitat重排任务上的实验表明:第一,该方法在更复杂的重排任务上优于专用模仿基线,且接近使用行为克隆技能的最优规划器基准,而无需任何最优任务计划或技能标注的完整任务演示;第二,消融实验证明,可靠的评论家引导候选选择对多模态行为协调至关重要;第三,扩展实验显示,该方法能处理更大行为库,在链式目标延长任务时仍保持强于专用模仿基线的性能。结果表明,显式技能抽象并非长时序重排的必要前提,隐式行为协调为数据驱动的替代方案提供了前景。

原文摘要 · Abstract (English)

Long-horizon robotic rearrangement tasks are often treated as skill sequencing problems, requiring predefined skills, skill labels, or boundaries, and task-specific switching logic. Although effective, such explicit skill abstractions can become difficult to scale as the number of behaviors and the task horizon increase. We instead formulate rearrangement as implicit-behavior coordination from unlabeled sub-task demonstrations, where skill-like behaviors are learned directly from mixed behavior data and coordinated through value-guided action selection. Experiments in Habitat rearrangement tasks support this formulation in three ways. First, our method outperforms task-specific imitation baselines on more complex rearrangement tasks and approaches an oracle-planner baseline with behavior-cloned skills, while using no oracle task plan or skill-labeled full-task demonstrations. Second, ablations show that reliable critic-guided candidate selection is essential for coordinating multi-modal behaviors. Third, scaling experiments show that the method handles larger behavior repertoires and maintains stronger performance than task-specific imitation baselines as chained targets extend the horizon. These results suggest that explicit skill abstraction is not a prerequisite for long-horizon rearrangement, and that implicit-behavior coordination offers a promising data-driven alternative to explicit skill-based pipelines.

机器人重排任务隐式协调无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。