arXiv:2508.02644cs.AI2025-08AAAI被引 5

用分散损失提升机器人操控中细微差异的识别能力

D2PPO: Diffusion Policy Policy Optimization with Dispersive Loss

  • 在扩散策略中引入分散损失,让相似观察映射为可区分特征
  • 复杂任务上微调后性能提升26.1%,达到新SOTA
  • 适合需要精细动作控制的机器人学习场景

扩散策略在高维空间中自然建模多模态动作分布,擅长机器人操作。然而,扩散策略存在表示坍缩问题:语义相似的观测被映射为无法区分的特征,削弱了对复杂操作中细微但关键差异的处理能力。为此,我们提出D2PPO(带分散损失的扩散策略策略优化),通过将每批隐藏表示全部视为负样本对,施加分散损失正则化,迫使网络学习相似观测的判别性表示,从而识别精确操作所需的细微差异。评估显示,早期层正则化利于简单任务,晚期层正则化显著提升复杂任务表现。在RoboMimic基准上,预训练平均提升22.7%,微调后提升26.1%,达到新SOTA。与最先进方法相比,弗兰卡·埃米卡·熊猫机器人真实实验中成功率显著更高,尤其在复杂任务中优势明显。

原文摘要 · Abstract (English)

Diffusion policies excel at robotic manipulation by naturally modeling multimodal action distributions in high-dimensional spaces. Nevertheless, diffusion policies suffer from diffusion representation collapse: semantically similar observations are mapped to indistinguishable features, ultimately impairing their ability to handle subtle but critical variations required for complex robotic manipulation. To address this problem, we propose D2PPO (Diffusion Policy Policy Optimization with Dispersive Loss). D2PPO introduces dispersive loss regularization that combats representation collapse by treating all hidden representations within each batch as negative pairs. D2PPO compels the network to learn discriminative representations of similar observations, thereby enabling the policy to identify subtle yet crucial differences necessary for precise manipulation. In evaluation, we find that early-layer regularization benefits simple tasks, while late-layer regularization sharply enhances performance on complex manipulation tasks. On RoboMimic benchmarks, D2PPO achieves an average improvement of 22.7% in pre-training and 26.1% after fine-tuning, setting new SOTA results. In comparison with SOTA, results of real-world experiments on a Franka Emika Panda robot show the excitingly high success rate of our method. The superiority of our method is especially evident in complex tasks. Project page: https://guowei-zou.github.io/d2ppo/

机器人操控扩散模型策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。