arXiv:2509.14848cs.LGeess.SP2025-09

用多精度模拟器节省训练成本,智能选择最值得投入的仿真级别。

Multi-Fidelity Hybrid Reinforcement Learning via Information Gain Maximization

  • 基于信息增益动态选择不同精度的仿真环境进行训练。
  • 在固定预算下,性能优于现有混合强化学习方法。
  • 适合资源受限但有多套仿真系统的实际场景使用。

强化学习策略优化通常需要与高精度环境模拟器进行大量交互,成本高昂或难以实现。离线强化学习虽可利用预收集数据训练,但效果受数据规模与质量限制。混合离线-在线强化学习结合了离线数据与单个模拟器的交互,但在真实场景中常存在多个精度与计算成本各异的模拟器。本文研究在固定成本预算下的多精度混合强化学习,提出基于信息增益最大化的多精度混合强化学习(MF-HRL-IGM)算法,通过自举法实现基于信息增益的仿真精度选择。理论分析证明该算法具备无遗憾性质,实验表明其在性能上显著优于现有基准方法。

原文摘要 · Abstract (English)

Optimizing a reinforcement learning (RL) policy typically requires extensive interactions with a high-fidelity simulator of the environment, which are often costly or impractical. Offline RL addresses this problem by allowing training from pre-collected data, but its effectiveness is strongly constrained by the size and quality of the dataset. Hybrid offline-online RL leverages both offline data and interactions with a single simulator of the environment. In many real-world scenarios, however, multiple simulators with varying levels of fidelity and computational cost are available. In this work, we study multi-fidelity hybrid RL for policy optimization under a fixed cost budget. We introduce multi-fidelity hybrid RL via information gain maximization (MF-HRL-IGM), a hybrid offline-online RL algorithm that implements fidelity selection based on information gain maximization through a bootstrapping approach. Theoretical analysis establishes the no-regret property of MF-HRL-IGM, while empirical evaluations demonstrate its superior performance compared to existing benchmarks.

强化学习多精度信息增益策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。