arXiv:2505.23003cs.LGcs.AI2025-05中稿 · KDD被引 3

用模拟器补足少量离线数据,提升鲁棒强化学习的样本效率。

Hybrid Cross-domain Robust Reinforcement Learning

  • 结合离线数据与在线模拟器,通过不确定性过滤选择可靠样本。
  • 在多种任务上超越现有方法,显著减少对大量离线数据的依赖。
  • 适合数据稀缺但需高鲁棒性的实际应用场景,如机器人控制。

鲁棒强化学习旨在学习在环境不确定性下仍有效的策略,这在真实场景中因动态变化而常见。现有离线鲁棒RL方法需大量数据,收集成本高。使用不完美模拟器可加速数据采集,但存在动态偏差问题。本文提出首个混合跨域鲁棒强化学习框架HYDRO,利用在线模拟器补充有限的离线数据。通过测量并最小化模拟器与不确定性集中最差模型之间的性能差距,HYDRO采用新颖的不确定性过滤和优先采样策略,选取最相关且可靠的模拟样本。大量实验表明,HYDRO在多个任务上均优于现有方法,证明其在提升离线鲁棒RL样本效率方面的潜力。

原文摘要 · Abstract (English)

Robust reinforcement learning (RL) aims to learn policies that remain effective despite uncertainties in its environment, which frequently arise in real-world applications due to variations in environment dynamics. The robust RL methods learn a robust policy by maximizing value under the worst-case models within a predefined uncertainty set. Offline robust RL algorithms are particularly promising in scenarios where only a fixed dataset is available and new data cannot be collected. However, these approaches often require extensive offline data, and gathering such datasets for specific tasks in specific environments can be both costly and time-consuming. Using an imperfect simulator offers a faster, cheaper, and safer way to collect data for training, but it can suffer from dynamics mismatch. In this paper, we introduce HYDRO, the first Hybrid Cross-Domain Robust RL framework designed to address these challenges. HYDRO utilizes an online simulator to complement the limited amount of offline datasets in the non-trivial context of robust RL. By measuring and minimizing performance gaps between the simulator and the worst-case models in the uncertainty set, HYDRO employs novel uncertainty filtering and prioritized sampling to select the most relevant and reliable simulator samples. Our extensive experiments demonstrate HYDRO's superior performance over existing methods across various tasks, underscoring its potential to improve sample efficiency in offline robust RL.

强化学习鲁棒性样本效率模拟器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。