揭示贝尔曼方程的三重对偶性,统一强化学习与控制理论方法
Generalised Bellman recurrence and three dualities in sequential decision-making
- 基于充分统计量、回报递归和不确定性聚合的一致性推导贝尔曼方程
- 三个对偶关系(概率-回报、回报-聚合、聚合-概率)源自同一构造
- 为强化学习、控制与决策理论提供统一数学框架
贝尔曼方程的形式由三个条件共同决定:动态过程可通过充分统计量分解,回报可递归分解,且不确定性聚合与二者相容。当三者在共同状态上同时成立时,贝尔曼方程自然出现;若任一条件失效,可通过扩展状态空间或变形回报/动态恢复可解性。这三个条件还导出三种对偶性:概率与回报之间、回报与聚合之间、聚合与概率之间的对偶。本框架将强化学习、控制与决策理论中独立发展的方法统一于同一构造之下。
原文摘要 · Abstract (English)
What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: that the dynamics decomposes through sufficient statistics, that the return decomposes recursively, and that the aggregation of uncertainty is compatible with both. When all three conditions hold on a common state, the Bellman equation arises from their mutual consistency; when one fails, tractability can often be recovered by augmenting the state or by deforming return or dynamics. The same conditions are shown to give rise to three dualities: one between probability and return, one between return and aggregation, and one between aggregation and probability. Our framework reveals these dualities as arising from a single construction, unifying methods developed separately across reinforcement learning, control, and decision theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。