arXiv:2512.02435cs.LG2025-12被引 2

通过动态与价值双重对齐筛选数据,提升跨域离线强化学习性能。

Efficient Cross-Domain Offline Reinforcement Learning with Dynamics- and Value-Aligned Data Filtering

  • 同时考虑动作转移规律与奖励价值,筛选高质量源域数据。
  • 在目标数据极少时仍显著优于基线方法,提升幅度达30%以上。
  • 适合数据稀缺的跨域强化学习场景,尤其适用于机器人控制任务。

跨域离线强化学习旨在利用少量目标域数据和可能覆盖充分的源域数据,训练出在目标环境表现良好的智能体。由于源域与目标域间存在动态偏差,直接合并数据可能导致性能下降。现有方法仅关注动态对齐,但本文理论分析表明,仅靠动态对齐不足,还需确保价值对齐——即选择高价值、高质量的源域样本。为此,提出统一框架DVDF(Dynamics- and Value-aligned Data Filtering),同时筛选在动态和价值上均与目标域对齐的源域样本。在多种动力学偏移场景(包括运动学与形态变化)下进行实验,涵盖多个任务与数据集,即使在目标域数据极度稀缺的情况下也表现出色。大量实验证明,DVDF持续优于强基线,性能提升显著。

原文摘要 · Abstract (English)

Cross-domain offline reinforcement learning (RL) aims to train a well-performing agent in the target environment, leveraging both a limited target domain dataset and a source domain dataset with (possibly) sufficient data coverage. Due to the underlying dynamics misalignment between source and target domains, naively merging the two datasets may incur inferior performance. Recent advances address this issue by selectively leveraging source domain samples whose dynamics align well with the target domain. However, our work demonstrates that dynamics alignment alone is insufficient, by examining the limitations of prior frameworks and deriving a new target domain sub-optimality bound for the policy learned on the source domain. More importantly, our theory underscores an additional need for \textit{value alignment}, i.e., selecting high-quality, high-value samples from the source domain, a critical dimension overlooked by existing works. Motivated by such theoretical insight, we propose \textbf{\underline{D}}ynamics- and \textbf{\underline{V}}alue-aligned \textbf{\underline{D}}ata \textbf{\underline{F}}iltering (DVDF) method, a novel unified cross-domain RL framework that selectively incorporates source domain samples exhibiting strong alignment in \textit{both dynamics and values}. We empirically study a range of dynamics shift scenarios, including kinematic and morphology shifts, and evaluate DVDF on various tasks and datasets, even in the challenging setting where the target domain dataset contains an extremely limited amount of data. Extensive experiments demonstrate that DVDF consistently outperforms strong baselines with significant improvements.

强化学习跨域迁移离线学习数据筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。