挑战强化学习中价值函数的必要性,提出应重新审视其建模假设。
Is there Value in Reinforcement Learning?
- 指出政策梯度方法并非真正无价值,仍需价值表征来学习。
- 强调标准强化学习假设(如风险中性、马尔可夫环境)影响价值定义。
- 主张认知模型应考虑计算复杂度而非仅参数统计复杂度。
行动价值在主流强化学习行为模型中扮演核心角色,但其是否被显式表征长期存在争议。批评者认为应偏好策略梯度(PG)模型以规避此问题,但我们指出该方案并不成立:尽管PG方法不依赖显式价值进行决策(刺激-反应映射),却仍需价值表征用于学习。因此,单纯转向PG并不能真正消除价值。更根本的问题在于,价值需求源于标准强化学习框架的底层假设,而非具体算法选择。以往研究多默认这些假设,仅争论优化方法(如PG vs VB)。我们主张将讨论焦点转向批判性评估这些建模假设本身。尤其从实验角度,当放松标准假设(如风险中性、完全可观测、马尔可夫环境、指数折扣)时,价值概念需重新思考,这在自然环境中更可能。最后,我们以价值争论为例,倡导认知科学中应采用更精细的‘算法视角’而非纯统计视角来定义模型。分析表明,模型复杂度除参数统计复杂度外,还应包括计算复杂度。
原文摘要 · Abstract (English)
Action-values play a central role in popular Reinforcement Learing (RL) models of behavior. Yet, the idea that action-values are explicitly represented has been extensively debated. Critics had therefore repeatedly suggested that policy-gradient (PG) models should be favored over value-based (VB) ones, as a potential solution for this dilemma. Here we argue that this solution is unsatisfying. This is because PG methods are not, in fact, "Value-free" -- while they do not rely on an explicit representation of Value for acting (stimulus-response mapping), they do require it for learning. Hence, switching to PG models is, per se, insufficient for eliminating Value from models of behavior. More broadly, the requirement for a representation of Value stems from the underlying assumptions regarding the optimization objective posed by the standard RL framework, not from the particular algorithm chosen to solve it. Previous studies mostly took these standard RL assumptions for granted, as part of their conceptualization or problem modeling, while debating the different methods used to optimize it (i.e., PG or VB). We propose that, instead, the focus of the debate should shift to critically evaluating the underlying modeling assumptions. Such evaluation is particularly important from an experimental perspective. Indeed, the very notion of Value must be reconsidered when standard assumptions (e.g., risk neutrality, full-observability, Markovian environment, exponential discounting) are relaxed, as is likely in natural settings. Finally, we use the Value debate as a case study to argue in favor of a more nuanced, algorithmic rather than statistical, view of what constitutes "a model" in cognitive sciences. Our analysis suggests that besides "parametric" statistical complexity, additional aspects such as computational complexity must also be taken into account when evaluating model complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。