arXiv:2511.04671cs.ROcs.AI2025-11被引 3

用人类动作的噪声版本训练机器人,让机器人学会人类意图而不照搬不切实际的动作。

X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations

  • 将人类动作视为机器人动作的噪声版本,在扩散模型中仅在高噪声阶段使用人类示范。
  • 在5个真实任务中提升成功率16%,优于直接混合训练和人工筛选数据。
  • 适合想低成本获取人类示范并高效迁移至机器人的研究者或工程师。

人类视频是机器人学习中可扩展的训练数据来源。然而,人与机器人在身体形态上差异显著,许多人类动作无法直接在机器人上执行。尽管如此,这些示范仍包含丰富的物体交互线索和任务意图。我们的目标是从这种粗略引导中学习,而不传递依赖具体身体形态的不可行执行策略。受生成建模中低质量数据处理方法的启发,特别是环境扩散(Ambient Diffusion)在高噪声阶段引入低质量数据的思想,我们提出X-Diffusion:一种基于扩散模型的跨体感学习框架。该方法将人类动作视为机器人动作的噪声形式,在前向扩散过程中随着噪声增加,体感差异逐渐消失,而任务相关指导得以保留。通过选择性地在高噪声阶段训练扩散策略,实现对易获取的人类视频的有效利用,同时保证机器人执行可行性。在五个真实世界操作任务中,X-Diffusion相比朴素联合训练和人工数据过滤,平均成功率提升16%。

原文摘要 · Abstract (English)

Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making many human actions infeasible for direct execution on a robot. Still, these demonstrations convey rich object-interaction cues and task intent. Our goal is to learn from this coarse guidance without transferring embodiment-specific, infeasible execution strategies. Recent advances in generative modeling tackle a related problem of learning from low-quality data. In particular, Ambient Diffusion is a recent method for diffusion modeling that incorporates low-quality data only at high-noise timesteps of the forward diffusion process. Our key insight is to view human actions as noisy counterparts of robot actions. As noise increases along the forward diffusion process, embodiment-specific differences fade away while task-relevant guidance is preserved. Based on these observations, we present X-Diffusion, a cross-embodiment learning framework based on Ambient Diffusion that selectively trains diffusion policies on noised human actions. This enables effective use of easy-to-collect human videos without sacrificing robot feasibility. Across five real-world manipulation tasks, we show that X-Diffusion improves average success rates by 16% over naive co-training and manual data filtering. The project website is available at https://portal-cornell.github.io/X-Diffusion/.

扩散模型机器人学习人类示范跨体感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。