改进强化学习中价值感知模型的训练,让预测更准确可靠
Calibrated Value-Aware Model Learning with Probabilistic Environment Models
- 提出校准价值感知损失的新方法,解决原有方法不准确的问题
- 证明确定性模型可准确预测价值,但随机模型仍具优势
- 揭示损失校准与模型结构、辅助损失间的相互作用机制
价值感知模型学习(value-aware model learning)强调模型需生成准确的价值估计,已在基于模型的强化学习中受到关注。已有研究采用MuZero损失,通过惩罚模型对价值函数的预测误差来优化性能,但其理论分析尚不充分。本文系统分析了包含MuZero损失在内的价值感知损失家族,发现这些损失在常规使用下为未校准的代理损失,无法始终恢复正确的模型与价值函数。基于此,我们提出修正方案以解决该问题。此外,我们研究了损失校准性、潜在模型架构及常用辅助损失之间的相互影响。结果表明,尽管确定性模型足以预测准确价值,但学习校准的随机模型仍具优势。
原文摘要 · Abstract (English)
The idea of value-aware model learning, that models should produce accurate value estimates, has gained prominence in model-based reinforcement learning. The MuZero loss, which penalizes a model's value function prediction compared to the ground-truth value function, has been utilized in several prominent empirical works in the literature. However, theoretical investigation into its strengths and weaknesses is limited. In this paper, we analyze the family of value-aware model learning losses, which includes the popular MuZero loss. We show that these losses, as normally used, are uncalibrated surrogate losses, which means that they do not always recover the correct model and value function. Building on this insight, we propose corrections to solve this issue. Furthermore, we investigate the interplay between the loss calibration, latent model architectures, and auxiliary losses that are commonly employed when training MuZero-style agents. We show that while deterministic models can be sufficient to predict accurate values, learning calibrated stochastic models is still advantageous.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。