让机器人学会语义对齐的原子技能,实现多任务灵活操作
Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
- 通过语义对比学习分解演示,构建可复用的原子技能库
- 预测关键姿态实现技能平滑衔接,长程执行成功率提升37%
- 适合需要多任务泛化能力的机器人操控场景
将模仿学习扩展到多样化的多任务机器人操作仍面临挑战,主要源于演示质量不佳、行为多模态性以及任务间破坏性干扰。现有基于技能的方法常生成受语言结构偏倚或跨任务语义不一致的技能,限制泛化能力。本文提出AtomSkill框架,从演示中学习语义对齐的原子技能空间,并通过关键姿态想象实现鲁棒的长时序执行。方法包括:(1) 语义对比技能对齐,将演示划分为可变长度的原子技能,采用对比目标同时保证语义一致性和时间连贯性,形成紧凑可复用的技能库;(2) 关键姿态预测驱动的动作解码,策略同时预测技能终止关键姿态与即时动作,支持进度感知的技能切换。推理阶段,原子技能扩散采样器生成合理技能序列,预测的关键姿态自动触发平滑技能链式执行。仿真与真实世界实验表明,AtomSkill持续优于当前最先进的模仿学习与技能基基线。项目页:https://atom-skill.github.io。
原文摘要 · Abstract (English)
Scaling imitation learning to diverse multi-task robot manipulation remains challenging due to suboptimal demonstrations, behavioral multi-modality, and destructive interference across tasks. While skill-based methods offer a promising direction by decomposing behaviors into reusable abstractions, existing approaches often learn skills that are either biased toward linguistic structure or lack semantic alignment across tasks, limiting generalization. In this work, we propose AtomSkill, a novel framework that learns a semantically aligned Atomic Skill Space from demonstrations and enables robust long-horizon execution through keypose imagination. Our method introduces: (1) semantic contrastive skill alignment, which partitions demonstrations into variable-length atomic skills and employs a contrastive objective to jointly enforce semantic consistency and temporal coherence, yielding a compact and reusable skill library; and (2) action decoding with keypose imagining, where the policy predicts both a skill's terminal keypose and immediate actions, thereby supporting progress-aware skill transitions. During inference, an atomic skill diffusion sampler generates plausible skill sequences, while predicted keyposes autonomously trigger smooth skill chaining. Extensive experiments in simulation and real-world settings show that AtomSkill consistently outperforms state-of-the-art imitation learning and skill-based baselines. Project page: https://atom-skill.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。