提出抗环境误差的策略评估数据采集方法,降低真实场景评估方差。
Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation
- 双循环梯度法优化数据采集策略,兼顾效率与鲁棒性。
- 在模拟到真实迁移中,方差显著低于现有方法。
- 适合对评估稳定性要求高的强化学习应用。
强化学习中的策略评估常因在线评估方差过高而受限。已有行为策略搜索方法虽能降低方差,但未考虑转移函数不确定性。实际中,仿真环境因建模误差与近似限制,其转移模型与真实世界存在差异,导致仿真训练的行为策略在真实环境部署时仍产生高方差,依赖昂贵的真实样本评估。本文提出一种基于双循环梯度的算法,用于学习既高效又对转移不确定性鲁棒的行为策略。理论上,推导出新型转移-方差梯度表达式,并建立全局收敛性保证;数值上,验证所提方法对转移扰动的敏感性显著低于现有方法,支持其实际应用价值。
原文摘要 · Abstract (English)
In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been proposed to learn data-collecting policies tailored to reduce online evaluation variance. However, these approaches do not account for uncertainties in the transition functions. In practice, simulator transitions often differ from the real world due to modeling errors or approximation limitations. As a result, behavior policies trained in simulation may still yield high variance when deployed in real environments, leading to costly reliance on real-world evaluation samples. In this work, we propose a double-loop gradient-based algorithm for learning behavior policies that are both efficient and robust to transition uncertainty. Theoretically, we derive novel transition-variance gradient expressions and establish global convergence guarantees for the algorithm. Numerically, we demonstrate that our method is less sensitive to transition perturbations than existing approaches, providing supportive evidence for its practical utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。