用可组合技能重构任务,实现机器人零样本跨任务操作
Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation

- 将演示分解为原子技能-动作对,支持可组合推理
- 在AGNOSTOS上实现零样本跨任务泛化,成功率显著提升
- 适合需要灵活迁移技能的机器人系统研究者
开放世界机器人操作中的跨任务泛化是核心挑战,关键在于从已见任务中提取可迁移的操作知识。现有上下文学习方法仅提供低层连续动作序列作为上下文,无法捕捉可组合的技能知识,导致模型退化为表面轨迹模仿。我们提出「分解与重组」技能推理框架,以原子技能-动作对作为中间表示。该方法将已见演示分解为可解释的技能-动作对齐,使模型能通过组合推理重构未知任务的技能。具体地,我们结合视觉-语义检索与规划器生成的技能序列,构建任务自适应动态演示库,并辅以覆盖感知的静态库填补缺失技能模式。两者共同生成包含完整技能信息的演示,显式激发组合推理能力以完成技能编排与执行顺序决策。在AGNOSTOS基准和真实环境中的实验验证了本方法的零样本跨任务泛化能力。
原文摘要 · Abstract (English)
Cross-task generalization is a core challenge in open-world robotic manipulation, and the key lies in extracting transferable manipulation knowledge from seen tasks. Recent in-context learning approaches leverage seen task demonstrations to generate actions for unseen tasks without parameter updates. However, existing methods provide only low-level continuous action sequences as context, failing to capture composable skill knowledge and causing models to degenerate into superficial trajectory imitation. We propose Decompose and Recompose, a skill reasoning framework using atomic skill-action pairs as intermediate representations. Our approach decomposes seen demonstrations into interpretable skill--action alignments, enabling the model to recompose these skills for unseen tasks through compositional reasoning. Specifically, we construct a task-adaptive dynamic demonstration library via visual-semantic retrieval combined with skill sequences from a planning agent, complemented by a coverage-aware static library to fill missing skill patterns. Together, these yield skill-comprehensive demonstrations that explicitly elicit compositional reasoning for skill composition and execution ordering. Experiments on the AGNOSTOS benchmark and real-world environments validate our method's zero-shot cross-task generalization capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。