arXiv:2606.08743cs.RO2026-06

用扩散模型挖掘机器人新行为,突破数据少时的创新瓶颈

Guided Discovery of New Behaviors using Diffusion Policies

论文配图:Guided Discovery of New Behaviors using Diffusion Policies
图 1 · 摘自论文原文
  • 结合费曼-卡茨校正与新型引导势能,系统引导采样向稀有行为靠近
  • 在多个操作环境中持续发现新可行动作,显著提升行为多样性
  • 适合研究机器人生成模型与行为探索的学者参考

扩散模型已成为机器人领域生成建模的强大工具,扩散策略擅长建模多模态动作轨迹分布。然而,在示范数据有限时,标准采样常重复主导行为而忽略有效但罕见的模式,限制了新解决方案的发现。现有方法如引导技术或结合强化学习的扩散模型,要么将样本推入不可行区域,要么难以跳出局部最优,无法系统性发掘多样化行为。为此,我们提出一种框架,结合费曼-卡茨校正器与新颖的引导势能,系统引导扩散策略采样趋向有前景但未被充分代表的样本。这些轨迹通过基于采样的轨迹优化进行精炼,并重新纳入训练集以重训扩散策略。该方法有效挖掘并修复新轨迹,实现多样化且可执行行为的系统性发现。我们在多个操纵环境中验证了该框架的有效性,均能持续发现新行为。

原文摘要 · Abstract (English)

Diffusion models have become a powerful tool for generative modeling in robotics, with diffusion policies excelling at modeling multimodal action-trajectory distributions. However, when demonstrations are limited, standard sampling often reproduces dominant behaviors while neglecting valid but rare modes, limiting the discovery of novel solutions. Existing approaches, such as guidance methods or combining reinforcement learning with diffusion, either push samples into infeasible regions or struggle to escape local minima, failing to systematically uncover diverse behaviors. To address these challenges, we propose a framework that combines Feynman-Kac correctors with a novel guiding potential that systematically guides diffusion policy samples towards promising yet underrepresented samples. These trajectories are refined using sampling-based trajectory optimization and reincorporated into the training set to retrain the diffusion policy. Our method effectively mines and repairs novel trajectories, enabling the systematic discovery of diverse and executable behaviors. We demonstrate the effectiveness of our framework across a range of manipulation environments, consistently discovering new behaviors.

机器人扩散模型行为发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。