arXiv:2503.02125cs.LG2025-03被引 2

提出简单有效的密度比估计算法,显著提升离策略评估准确性

Average-DICE: Stationary Distribution Correction by Regression

  • 用平均折扣重要性采样比估计密度比,计算高效且无偏
  • 在Mujoco环境上表现优于或媲美顶尖方法,部分场景提升数个数量级
  • 适合需要高精度离策略评估的研究者,尤其关注算法稳定性

离策略策略评估(OPE)长期受静态状态分布不匹配影响,导致估计不稳定且不准确。现有方法通过估计密度比修正分布偏移,但常依赖昂贵优化或基于反向贝尔曼更新,难以超越简单基线。本文提出AVG-DICE,一种基于蒙特卡洛的密度比估计器,通过平均折扣重要性采样比实现无偏且一致的校正。该方法可自然扩展至非线性函数逼近,采用回归方式进行近似,在Mujoco Gym环境的OPE任务中测试,并与最先进密度比估计器使用其报告超参数进行对比。实验表明,AVG-DICE至少达到当前最优水平,某些情况下性能提升达数个数量级。然而敏感性分析显示,最佳超参数随折扣因子变化显著,建议针对不同场景重新调参。

原文摘要 · Abstract (English)

Off-policy policy evaluation (OPE), an essential component of reinforcement learning, has long suffered from stationary state distribution mismatch, undermining both stability and accuracy of OPE estimates. While existing methods correct distribution shifts by estimating density ratios, they often rely on expensive optimization or backward Bellman-based updates and struggle to outperform simpler baselines. We introduce AVG-DICE, a computationally simple Monte Carlo estimator for the density ratio that averages discounted importance sampling ratios, providing an unbiased and consistent correction. AVG-DICE extends naturally to nonlinear function approximation using regression, which we roughly tune and test on OPE tasks based on Mujoco Gym environments and compare with state-of-the-art density-ratio estimators using their reported hyperparameters. In our experiments, AVG-DICE is at least as accurate as state-of-the-art estimators and sometimes offers orders-of-magnitude improvements. However, a sensitivity analysis shows that best-performing hyperparameters may vary substantially across different discount factors, so a re-tuning is suggested.

强化学习离策略评估密度比估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。