arXiv:2412.20537cs.LG2024-12

研究发现模型准确性再高,价值扩展方法的采样效率提升也有限。

Diminishing Return of Value Expansion Methods

  • 用理想模型消除误差后,滚动步长越长效率越高但增益递减。
  • 模型精度提升对效率改善微乎其微,不如直接用模型无关方法。
  • 适合关注模型基强化学习瓶颈的科研人员阅读。

基于模型的强化学习旨在提升采样效率,但动态模型的准确性与累积误差常被视为主要限制。本文通过使用理想动态模型(oracle dynamics models)消除累积误差,实证研究了改进模型对基于模型价值扩展方法采样效率的影响。结果表明:第一,更长的滚动预测步长虽能提升效率,但每增加一步带来的收益迅速衰减;第二,模型精度提升相比使用相同步长的训练模型,对采样效率的改进微不足道。这种效率增长的边际递减现象在与模型无关的价值扩展方法对比时尤为显著——后者性能相当,却无计算开销。因此,模型基方法的瓶颈不在于模型准确性。即使模型完美,也无法获得压倒性采样效率优势。这挑战了‘模型精度是主要制约’的普遍认知。

原文摘要 · Abstract (English)

Model-based reinforcement learning aims to increase sample efficiency, but the accuracy of dynamics models and the resulting compounding errors are often seen as key limitations. This paper empirically investigates potential sample efficiency gains from improved dynamics models in model-based value expansion methods. Our study reveals two key findings when using oracle dynamics models to eliminate compounding errors. First, longer rollout horizons enhance sample efficiency, but the improvements quickly diminish with each additional expansion step. Second, increased model accuracy only marginally improves sample efficiency compared to learned models with identical horizons. These diminishing returns in sample efficiency are particularly noteworthy when compared to model-free value expansion methods. These model-free algorithms achieve comparable performance without the computational overhead. Our results suggest that the limitation of model-based value expansion methods cannot be attributed to model accuracy. Although higher accuracy is beneficial, even perfect models do not provide unrivaled sample efficiency. Therefore, the bottleneck exists elsewhere. These results challenge the common assumption that model accuracy is the primary constraint in model-based reinforcement learning.

强化学习采样效率模型基价值扩展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。