用扩散薛定谔桥对齐跨域动态,实现无目标环境交互的迁移强化学习
Bridging Dynamics Gaps via Diffusion Schrödinger Bridge for Cross-Domain Reinforcement Learning
- 通过扩散薛定谔桥将源域转移分布与目标域离线演示对齐
- 在源域内完成策略学习,实测性能超越现有最佳方法
- 适合缺乏目标环境访问和奖励信号的迁移学习场景
跨域强化学习旨在应对源域与目标域间动态变化时学习可迁移策略。其核心挑战在于无法获取目标域环境交互及奖励监督,导致无法直接进行策略学习。为此,本文提出一种新框架BDGxRL,利用扩散薛定谔桥(DSB)将源域转移分布与目标域动态(由离线演示编码)对齐。同时引入奖励调制机制,基于状态转移估计奖励,并应用于经DSB对齐的样本,确保奖励与目标域动态的一致性。整个策略学习过程完全在源域内完成,无需访问目标环境或其奖励信号。在MuJoCo跨域基准上的实验表明,BDGxRL显著优于现有最优基线,在转移动态变化下展现出强适应性。
原文摘要 · Abstract (English)
Cross-domain reinforcement learning (RL) aims to learn transferable policies under dynamics shifts between source and target domains. A key challenge lies in the lack of target-domain environment interaction and reward supervision, which prevents direct policy learning. To address this challenge, we propose Bridging Dynamics Gaps for Cross-Domain Reinforcement Learning (BDGxRL), a novel framework that leverages Diffusion Schrödinger Bridge (DSB) to align source transitions with target-domain dynamics encoded in offline demonstrations. Moreover, we introduce a reward modulation mechanism that estimates rewards based on state transitions, applying to DSB-aligned samples to ensure consistency between rewards and target-domain dynamics. BDGxRL performs target-oriented policy learning entirely within the source domain, without access to the target environment or its rewards. Experiments on MuJoCo cross-domain benchmarks demonstrate that BDGxRL outperforms state-of-the-art baselines and shows strong adaptability under transition dynamics shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。