arXiv:2605.13054cs.LGcs.AI2026-05

用生成数据填补目标域数据不足,提升跨域离线强化学习性能

Bridging Domain Gaps with Target-Aligned Generation for Offline Reinforcement Learning

论文配图:Bridging Domain Gaps with Target-Aligned Generation for Offline Reinforcement Learning
图 1 · 摘自论文原文
  • 基于双分数模型生成与目标域一致的过渡数据
  • 在极少量目标域数据下仍显著优于现有方法
  • 适合数据稀缺场景下的跨域策略迁移

跨域离线强化学习旨在仅使用预先收集的数据集,将策略从源域迁移到目标域,其中环境动态可能不同。主要挑战在于如何利用源域数据的同时减少分布差异,尤其当目标域数据极度有限时。为此,我们提出目标对齐覆盖扩展(TCE)框架,通过理论分析指导源域数据的使用方式:直接引入靠近目标的转移数据,或通过目标对齐生成来扩展状态覆盖范围。TCE基于双分数生成模型,在扩展的状态区域内合成与目标域一致的转移数据。在多种跨域环境中的大量实验表明,TCE始终优于当前最先进的跨域离线强化学习基线方法。

原文摘要 · Abstract (English)

Cross-domain offline reinforcement learning aims to adapt a policy from a source domain to a target domain using only pre-collected datasets, where environment dynamics may differ. A key challenge is to leverage source data while reducing distributional mismatch, particularly when the target dataset is extremely limited. To address this, we propose Target-aligned Coverage Expansion (TCE), a framework that decides how source data should be used, either by directly incorporating target-near transitions or by expanding state coverage through target-aligned generation, guided by theoretical analysis. TCE builds on a dual score-based generative model to synthesize target-consistent transitions over an expanded state region. Extensive experiments across diverse cross-domain environments show that TCE consistently outperforms state-of-the-art cross-domain offline RL baselines.

离线RL跨域迁移生成模型数据扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。