arXiv:2602.02912cs.LGcs.AI2026-02

揭示强化学习中后验更新如何真实传递信息并决定行为激励

Notes on the Reward Representation of Posterior Updates

  • 用KL正则软更新实现贝叶斯后验,使决策更新具可解释的证据重加权机制
  • 后验更新仅决定相对激励信号,绝对奖励仍依赖上下文基线,无法唯一确定
  • 要求跨方向一致性引入新约束,统一不同条件顺序下的奖励描述

现代控制与强化学习中的许多思想将决策视为推理:从基准分布出发,在信号到达时进行更新。我们探讨这一过程能否从隐喻变为真实。研究发现,当一个KL正则化的软更新恰好对应于单一固定概率模型内的贝叶斯后验时,该更新变量即成为真实的信息传输通道。此时行为变化仅由该通道携带的证据驱动——更新必须可解释为对基准分布的证据重加权。由此得出明确识别结果:后验更新决定了相对的、上下文相关的激励信号以改变行为,但不唯一决定绝对奖励,后者仍可在上下文特定基线上保持模糊。若要求在不同更新方向间共享同一延续值,则进一步施加一致性约束,关联不同条件顺序对应的奖励描述。

原文摘要 · Abstract (English)

Many ideas in modern control and reinforcement learning treat decision-making as inference: start from a baseline distribution and update it when a signal arrives. We ask when this can be made literal rather than metaphorical. We study the special case where a KL-regularized soft update is exactly a Bayesian posterior inside a single fixed probabilistic model, so the update variable is a genuine channel through which information is transmitted. In this regime, behavioral change is driven only by evidence carried by that channel: the update must be explainable as an evidence reweighing of the baseline. This yields a sharp identification result: posterior updates determine the relative, context-dependent incentive signal that shifts behavior, but they do not uniquely determine absolute rewards, which remain ambiguous up to context-specific baselines. Requiring one reusable continuation value across different update directions adds a further coherence constraint linking the reward descriptions associated with different conditioning orders.

强化学习贝叶斯推理奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。