arXiv:2605.16871cs.RO2026-05

用基础模型生成带子目标的示范,让机器人长时序操作更可控。

SADP: Subgoal-Aware Diffusion Policy for Long-Horizon Manipulation Learned from Foundation Model Generated Demonstrations

论文配图:SADP: Subgoal-Aware Diffusion Policy for Long-Horizon Manipulation Learned from Foundation Model Generated Demonstrations
图 1 · 摘自论文原文
  • 基于基础模型自动生成带子目标标注的示范数据
  • 扩散策略同时接受任务和子目标描述,实现精准动作生成
  • 新增持续性评分头,支持在线子目标切换与进度监控

长时序机器人操作需要协调多个中间子目标并决定何时推进。然而,多数模仿学习方法仅在任务级示范上训练,未显式建模子目标或其执行进度。标准机器人学习数据集缺乏子目标层面的监督,使得显式子目标条件控制和在线转换建模难以实现。本文提出子目标感知扩散策略(SADP),利用基础模型自动生成带子目标标注的示范,并在此数据集上训练扩散策略。SADP通过显式自然语言子目标对动作生成进行条件化,同时引入轻量辅助头预测延续性得分,驱动在线子目标切换并支持阶段级进度监控。在RLBench仿真和真实世界中使用UR5e机器人的实验表明,SADP在保持竞争性任务性能的同时,揭示了对齐时间的子目标级执行信号,可用于进度监控。结果表明,显式子目标进展可融入扩散策略而不会降低任务级表现。

原文摘要 · Abstract (English)

Long-horizon robot manipulation requires policies to coordinate multiple intermediate subgoals and determine when to advance between them. However, most imitation learning methods are trained solely on task-level demonstrations, without explicitly modeling the active subgoal or its execution progress. This limitation is further exacerbated by the scarcity of subgoal-level supervision in standard robot learning datasets, which makes explicit subgoal-conditioned control and online transition modeling difficult to learn. To address this issue, this paper proposes Subgoal-Aware Diffusion Policy (SADP), a framework that leverages foundation models to autonomously generate subgoal-annotated demonstrations and trains diffusion policies on these datasets. SADP structures policy execution around explicit natural-language subgoals by conditioning action generation on both task-level and subgoal-level descriptions. A lightweight auxiliary head further predicts a continuation score that drives online subgoal switching and supports stage-level progress monitoring. Experiments in RLBench simulations and real-world evaluations on a UR5e robot demonstrate that SADP maintains competitive task performance while exposing temporally aligned subgoal-level execution signals for progress monitoring. These results show that explicit subgoal progression can be incorporated into a diffusion policy without degrading task-level performance.

扩散模型机器人操控子目标模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。