arXiv:2603.10980cs.RO2026-03中稿 · ICRA被引 2

用轻量级引导机制提升扩散策略的机器人操作鲁棒性

PPGuide: Steering Diffusion Policies with Performance Predictive Guidance

  • 通过自监督注意力学习自动识别成功/失败的关键动作片段
  • 在多个任务上实现稳定性能提升,避免错误累积导致失败
  • 适合需要实时纠错的机器人控制场景,无需额外训练数据

扩散策略在学习复杂、多模态的机器人操作行为方面表现出色。然而,生成的动作序列中的错误会随时间累积,可能导致失败。现有方法通过引入专家示范或学习预测世界模型来缓解问题,但可能计算开销大。本文提出性能预测引导(PPGuide),一种轻量级、基于分类器的推理时引导框架,可使预训练扩散策略避开失败模式。该方法采用新型自监督流程:利用注意力机制的多实例学习,自动识别策略回放中与成功或失败相关的观测-动作片段,并以此构建自标注数据训练性能预测器。推理时,该预测器提供实时梯度,引导策略选择更鲁棒的动作。我们在Robomimic和MimicGen基准的多样化任务上验证了PPGuide,均取得一致性能提升。

原文摘要 · Abstract (English)

Diffusion policies have shown to be very efficient at learning complex, multi-modal behaviors for robotic manipulation. However, errors in generated action sequences can compound over time which can potentially lead to failure. Some approaches mitigate this by augmenting datasets with expert demonstrations or learning predictive world models which might be computationally expensive. We introduce Performance Predictive Guidance (PPGuide), a lightweight, classifier-based framework that steers a pre-trained diffusion policy away from failure modes at inference time. PPGuide makes use of a novel self-supervised process: it uses attention-based multiple instance learning to automatically estimate which observation-action chunks from the policy's rollouts are relevant to success or failure. We then train a performance predictor on this self-labeled data. During inference, this predictor provides a real-time gradient to guide the policy toward more robust actions. We validated our proposed PPGuide across a diverse set of tasks from the Robomimic and MimicGen benchmarks, demonstrating consistent improvements in performance.

扩散模型机器人控制强化学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。