arXiv:2512.17086cs.AI2025-12被引 4

将通用人工智能的效用函数扩展至更广范畴,解决不确定性下的价值计算问题。

Value Under Ignorance in Universal Artificial Intelligence

  • 用模糊概率理论中的Choquet积分处理信念分布的不精确性
  • 揭示了在死亡解释下期望效用不可由Choquet积分刻画
  • 为强化学习中不确定环境的价值计算提供新视角

我们将强化学习中的通用智能体AIXI推广至更广泛的效用函数。为每个可能的交互历史分配效用时,需面对一个难题:某些假设仅预测历史的有限前缀,这常被解释为存在等于半测度损失的死亡概率。这一死亡解释提示了一种为历史前缀赋值的方法。我们主张,将信念分布视为不精确概率分布,半测度损失即为总无知量,更具自然性。这促使我们考虑使用不精确概率论中的Choquet积分计算期望效用,包括其可计算性分析。标准递归价值函数成为特例。然而,在死亡解释下最一般的期望效用无法被表征为此类Choquet积分。

原文摘要 · Abstract (English)

We generalize the AIXI reinforcement learning agent to admit a wider class of utility functions. Assigning a utility to each possible interaction history forces us to confront the ambiguity that some hypotheses in the agent's belief distribution only predict a finite prefix of the history, which is sometimes interpreted as implying a chance of death equal to a quantity called the semimeasure loss. This death interpretation suggests one way to assign utilities to such history prefixes. We argue that it is as natural to view the belief distributions as imprecise probability distributions, with the semimeasure loss as total ignorance. This motivates us to consider the consequences of computing expected utilities with Choquet integrals from imprecise probability theory, including an investigation of their computability level. We recover the standard recursive value function as a special case. However, our most general expected utilities under the death interpretation cannot be characterized as such Choquet integrals.

通用智能效用函数不精确概率强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。