用扩散模型统一学习多模态机器人轨迹,提升力控操作的鲁棒性。
Multimodal Diffusion Forcing for Forceful Manipulation
- 通过随机掩码训练扩散模型重建完整轨迹,捕捉模态间时序与跨模态依赖。
- 在模拟与真实环境中均实现强性能与噪声观测下的稳定表现。
- 适合需要多传感器融合和力控精准性的机器人操作研究者。
给定专家轨迹数据集,传统模仿学习通常建立从观测(如RGB图像)到动作的直接映射。然而,这类方法常忽略感官输入、动作与奖励之间的丰富交互关系,而这对于建模机器人行为和理解任务结果至关重要。本文提出多模态扩散强制(Multimodal Diffusion Forcing, MDF),一种统一的学习框架,用于从多模态机器人轨迹中学习,超越单纯动作生成。MDF不建模固定分布,而是通过随机部分掩码,并训练扩散模型以重构轨迹。该训练目标促使模型学习时间依赖性和跨模态依赖关系,例如预测动作对力信号的影响或从部分观测中推断状态。我们在模拟与真实世界环境中评估了MDF在接触丰富、力控要求高的操作任务上的表现。结果表明,MDF不仅具备多功能性,还在噪声观测下展现出优异性能与鲁棒性。
原文摘要 · Abstract (English)
Given a dataset of expert trajectories, standard imitation learning approaches typically learn a direct mapping from observations (e.g., RGB images) to actions. However, such methods often overlook the rich interplay between different modalities, i.e., sensory inputs, actions, and rewards, which is crucial for modeling robot behavior and understanding task outcomes. In this work, we propose Multimodal Diffusion Forcing, a unified framework for learning from multimodal robot trajectories that extends beyond action generation. Rather than modeling a fixed distribution, MDF applies random partial masking and trains a diffusion model to reconstruct the trajectory. This training objective encourages the model to learn temporal and cross-modal dependencies, such as predicting the effects of actions on force signals or inferring states from partial observations. We evaluate MDF on contact-rich, forceful manipulation tasks in simulated and real-world environments. Our results show that MDF not only delivers versatile functionalities, but also achieves strong performance, and robustness under noisy observations. More visualizations can be found on our $\href{https://unified-df.github.io}{website}$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。