arXiv:2602.12963cs.AI2026-02被引 1

最优策略能透露环境多少信息?答案是精确的 n log m 比特。

Calculating Mutual Information between a Reward Maximizer and its Environment

  • 通过信息论推导,最优策略与环境间互信息为 n log m 比特。
  • 无论短期、长期或平均奖励目标,该结果均成立。
  • 揭示了智能体实现最优行为所需隐含世界模型的最小信息量。

人工智能领域一个核心问题是:成功行为是否需要对世界的内部表征。本文量化了针对任意非恒定奖励函数的最优策略所携带的环境信息量。考虑一个具有 n 个状态和 m 个动作的受控马尔可夫过程(CMP),假设转移动态在所有可能中服从均匀先验。我们证明:观测到一个确定性最优策略时,其关于环境的互信息恰好为 n log m 比特。该结论适用于多种目标,包括有限时域、无限时域折扣及时间平均奖励最大化。研究给出了实现最优性所需的‘隐式世界模型’的信息下界。

原文摘要 · Abstract (English)

An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world. In this work, we quantify the amount of information an optimal policy provides about the underlying environment. We consider a Controlled Markov Process (CMP) with $n$ states and $m$ actions, assuming a uniform prior over the space of possible transition dynamics. We prove that observing a deterministic policy that is optimal for any non-constant reward function then conveys exactly $n \log m$ bits of information about the environment. Specifically, we show that the mutual information between the environment and the optimal policy is $n \log m$ bits. This bound holds across a broad class of objectives, including finite-horizon, infinite-horizon discounted, and time-averaged reward maximization. These findings provide a precise information-theoretic lower bound on the ``implicit world model'' necessary for optimality.

信息论强化学习最优策略世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。