arXiv:2512.02486cs.LG2025-12中稿 · ICLR被引 6

解决跨域离线强化学习在实际部署中对动态变化的脆弱性问题。

Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics Shifts

  • 提出双鲁棒的跨域贝尔曼算子,兼顾训练与测试阶段的稳定性
  • 在动态变化场景下性能超越强基线,尤其在目标域数据少时优势明显
  • 适合关注实际部署鲁棒性的离线强化学习研究者使用

单一领域离线强化学习常因数据覆盖不足而受限,跨域离线强化学习通过引入其他领域数据缓解此问题,但现有方法主要关注训练阶段对动态偏移的鲁棒性,忽视了实际部署时测试阶段对动态扰动的脆弱性。本文首次研究跨域离线强化学习在训练与测试双重场景下的鲁棒性。实证发现,跨域训练策略在评估时对动态扰动敏感,尤其当目标域数据稀少时表现更差。为此,提出新颖的鲁棒跨域贝尔曼(RCB)算子,在保持对分布外动态转移保守性的同时增强测试阶段鲁棒性,确保训练阶段鲁棒性。为进一步缓解RCB带来的值函数过估计或低估,引入动态值惩罚与Huber损失,构建实用的双鲁棒跨域离线强化学习(DROCO)算法。大量实验表明,DROCO在多种动态偏移场景下优于强基线,展现出更强的动态扰动鲁棒性。

原文摘要 · Abstract (English)

Single-domain offline reinforcement learning (RL) often suffers from limited data coverage, while cross-domain offline RL handles this issue by leveraging additional data from other domains with dynamics shifts. However, existing studies primarily focus on train-time robustness (handling dynamics shifts from training data), neglecting the test-time robustness against dynamics perturbations when deployed in practical scenarios. In this paper, we investigate dual (both train-time and test-time) robustness against dynamics shifts in cross-domain offline RL. We first empirically show that the policy trained with cross-domain offline RL exhibits fragility under dynamics perturbations during evaluation, particularly when target domain data is limited. To address this, we introduce a novel robust cross-domain Bellman (RCB) operator, which enhances test-time robustness against dynamics perturbations while staying conservative to the out-of-distribution dynamics transitions, thus guaranteeing the train-time robustness. To further counteract potential value overestimation or underestimation caused by the RCB operator, we introduce two techniques, the dynamic value penalty and the Huber loss, into our framework, resulting in the practical \textbf{D}ual-\textbf{RO}bust \textbf{C}ross-domain \textbf{O}ffline RL (DROCO) algorithm. Extensive empirical results across various dynamics shift scenarios show that DROCO outperforms strong baselines and exhibits enhanced robustness to dynamics perturbations.

强化学习离线学习鲁棒性跨域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。