arXiv:2607.10892cs.RO2026-07

一个扩散策略搞定多种积木推移,零样本跨域迁移。

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer

论文配图:A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer
图 1 · 摘自论文原文
  • 用强化学习从零训练单一扩散策略,支持多形状积木推移。
  • 在稀疏奖励仿真环境中成功学习多任务策略,零样本迁移到真实场景。
  • 适合需要泛化能力的机器人操控研究者,尤其关注现实部署。

扩散策略在基于行为克隆(BC)的机器人复杂动作学习中表现出色。本文探索使用强化学习(RL)从零开始训练扩散策略,用于多任务机器人操控。具体目标是训练一个单一扩散策略,完成多种形状积木的推移任务。所提框架采用简单的策略损失函数,即基于BC的重加权证据下界,可无缝集成到各类强化学习算法中。为解决无演示时的探索难题,引入逆向课程生成与目标中心表示。结合扩散策略的表达能力,该设计在稀疏奖励仿真环境下实现了多任务积木推移策略的学习。进一步评估表明,训练后的扩散策略在零样本条件下,能成功迁移到真实世界任务,适应不同目标位置、积木形状、重量及表面摩擦等环境变化,在测试条件下实现真实场景中的有效执行。

原文摘要 · Abstract (English)

Diffusion policies have shown promising empirical performance in representing and learning complex maneuvers for robots using behavior cloning (BC). In this paper, we explore training diffusion policies from scratch using reinforcement learning (RL) for multi-task robotic manipulation. Specifically, we aim to train a single diffusion policy for block-pushing tasks with multiple shapes. The proposed framework features a simple policy loss function, which is a reweighted evidence lower bound used in BC-based diffusion policy training and can seamlessly serve as the policy learning module in RL algorithms. To address the exploration challenges arising from the absence of demonstrations, we incorporate reverse curriculum generation and objective-centric representations. Combined with the expressiveness of diffusion policies, our design supports learning of multi-task block-pushing policies in our sparse-reward simulation setting. We further evaluate whether the trained diffusion policy transfers in zero-shot to real-world tasks under varying environmental conditions including goal positions, block shapes, block weights and surface friction, providing evidence that this pipeline can transfer to our real-world block-pushing setup under the tested variations.

扩散模型机器人操控零样本迁移强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。