从决策公理出发,证明软贝尔曼方程的合理性。
Rationalizing Boltzmann Rationality: An Axiomatic Characterization of Entropy-Regularized Policies
- 用独立性公理区分环境随机与选择随机,推导出软策略形式。
- 在选择节点上满足IIA和单调性时,唯一确定玻尔兹曼策略。
- 揭示软硬贝尔曼方程本质是代理是否重视自身选择权的设计取舍。
softmax策略 $π(a ackslashmid s) ackslashpropto ackslashexp(βQ(s,a))$ 是强化学习中默认的随机选择模型。尽管文献中已有鲁棒性、探索和优化等多重解释,但尚无一种从第一性原理唯一推导出该形式的论证。这导致一个基本矛盾:软贝尔曼方程中的熵奖励违反了支撑马尔可夫决策过程(MDP)奖励结构的独立性公理。本文通过区分两类随机性——偶然性与选择性,将冯·诺依曼-摩根斯坦(VNM)独立性仅限制于环境对基础前景的抽奖,证明在选择节点上对策略与价值函数施加独立于无关选项(IIA)和单调性,即可唯一确定玻尔兹曼策略、熵正则化表示及软贝尔曼方程。因此,软/硬贝尔曼方程的选择本质上是设计问题:代理是否重视自身的选择能力。本文进一步推导强化学习特有结论,如回报单调性及广义折扣下的收敛性,并整合经济学与信息论各自独立得出相同结构的研究成果,提供关于何时适用IIA的规范性评估。
原文摘要 · Abstract (English)
The softmax policy $π(a \mid s) \propto \exp(βQ(s,a))$ is the default model of stochastic choice in reinforcement learning (RL). Various justifications based on robustness, exploration, and optimization have been offered in the RL literature, but none uniquely derives the softmax form from first principles. This leaves a basic tension unresolved: the entropy bonus in the soft Bellman equation violates the Independence axiom that underwrites the Markov decision process (MDP) reward structure. We dissolve this tension by distinguishing two kinds of randomness: chance and choice. By restricting von Neumann-Morgenstern (VNM) Independence to environmental lotteries over base prospects, we show that imposing independence of irrelevant alternatives (IIA) and monotonicity on the policy and value functions at choice nodes uniquely determines the Boltzmann policy, the entropy-regularized representation, and the soft Bellman equation. The choice between the soft and hard Bellman equations thus reduces to a design decision: whether the agent values its own ability to choose. We develop RL-specific consequences, including return monotonicity and convergence under generalized discounting, and synthesize the independent lines from economics and information theory that arrive at the same structure, offering a normative assessment of when IIA is appropriate for agent design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。