用扩散模型直接编辑轨迹,跨域适配更灵活高效
xTED: Cross-Domain Adaptation via Diffusion-Based Trajectory Editing
- 设计扩散模型直接修改源域轨迹分布
- 在仿真与真实机器人上均显著提升跨域性能
- 无需定制模型,适配多种下游策略学习方法
复用不同领域的预收集数据是决策任务的可行方案,尤其在目标领域数据稀缺时。现有跨域策略迁移方法多聚焦于学习领域对应关系或修正项以促进策略学习,如任务/领域特定判别器、表示或策略,但常导致模型复杂或需特定建模,缺乏灵活性。本文提出跨域轨迹编辑框架xTED,采用专门设计的扩散模型实现跨域轨迹适配。该模型有效捕捉状态、动作、奖励间的复杂依赖及目标数据中的动态模式。通过添加噪声并利用预训练扩散模型去噪,源域轨迹可被转换为匹配目标域特性的形式,同时保留原始语义信息。此过程有效弥合底层域差距,提升源数据的状态真实性和动态可靠性,且可灵活集成于各类单域和跨域下游策略学习方法中。尽管结构简单,xTED在大量仿真与真实机器人实验中表现优异。
原文摘要 · Abstract (English)
Reusing pre-collected data from different domains is an appealing solution for decision-making tasks, especially when data in the target domain are limited. Existing cross-domain policy transfer methods mostly aim at learning domain correspondences or corrections to facilitate policy learning, such as learning task/domain-specific discriminators, representations, or policies. This design philosophy often results in heavy model architectures or task/domain-specific modeling, lacking flexibility. This reality makes us wonder: can we directly bridge the domain gaps universally at the data level, instead of relying on complex downstream cross-domain policy transfer procedures? In this study, we propose the Cross-Domain Trajectory EDiting (xTED) framework that employs a specially designed diffusion model for cross-domain trajectory adaptation. Our proposed model architecture effectively captures the intricate dependencies among states, actions, and rewards, as well as the dynamics patterns within target data. Edited by adding noises and denoising with the pre-trained diffusion model, source domain trajectories can be transformed to align with target domain properties while preserving original semantic information. This process effectively corrects underlying domain gaps, enhancing state realism and dynamics reliability in source data, and allowing flexible integration with various single-domain and cross-domain downstream policy learning methods. Despite its simplicity, xTED demonstrates superior performance in extensive simulation and real-robot experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。