arXiv:2606.12502physics.soc-phcs.AI2026-06被引 1

为目标导向智能体在资源约束下的价值建立数学理论,揭示其与信息类似的结构性规律。

A Mathematical Theory of Value: a synthesis on goal-directed agency under resource constraints

  • 将价值定义为资源转化为目标进展的速率,基于对数尺度建模。
  • 实测显示感知互信息与能力高度相关(Spearman ρ=0.977),且价值增量符合信息容量上限。
  • 适用于理解智能体对齐、资源配置与群体协同机制,适合研究智能体理性的学者。

我们提出价值——目标导向智能体创造、消耗与交换的量——是一种与信息同属一类的规律性结构量。借鉴香农方法,进行一次严格抽象:价值是智能体相对于其目标所设定的参考系,将资源转化为目标进展的速率。尺度不变性公理迫使采用对数度量形式 $V = \sum_i k_i \ln e_i$;复投资源的累积效应通过 Peters(2019)的遍历性论证也导出相同形式,作为一致性检验而非过度决定。我们推导出价值编码定理 $\Delta G \le I(X;Y)$;实现的价值可分解为 $G = D(q\|r) - D(q\|p)$。对群体而言,价值依赖于参考系,而价格则独立;一组共享资源并融合感知的智能体群,其价值上限为 $G_{\rm fleet} \le I(X;Y_{1:m}) \le H(X)$(新修正结果;此前的求和形式有误,已在 v5 中更正)。动力学层引出“是”与“应当”的不对称性,使对齐成为控制稳定性条件。我们在真实语言模型上预注册测试单参考系定律:感知互信息与实现能力高度一致(Spearman $ρ=0.977$,30 个模型×领域点);样本外 $\Delta G$ 与 $I(X;Y)$ 跨四种任务形状保持形状不变(n=42,斜率 0.953);过度自信可量化为耗散。后续预注册验证表明:耦合容量-区域预测——增长缺口律、联盟次模性与异或协同控制、联合上限、凯利选择——在真实智能体中于冻结区间内被证实;均场残差律 $\|Vg\|/γ$ 在各领域均未成立(群体无目标分散),已退回其纯数学范畴。贡献在于统一理论与随之而来的治理映射。

原文摘要 · Abstract (English)

We propose that value -- the quantity goal-directed agents create, destroy, and exchange -- is a lawful structural quantity in the same category as information. Following Shannon's method, we make one ruthless abstraction: value is the rate at which an agent converts a resource into goal-progress, relative to a frame fixed by its goal. A scale-invariance axiom forces a logarithmic measure, $V=\sum_i k_i\ln e_i$; compounding of a reinvested resource forces the same form via the ergodicity argument of Peters (2019) -- kin routes, a consistency check, not an over-determination. We derive a coding theorem of value, $ΔG \le I(X;Y)$; realized value decomposes as $G=D(q\|r)-D(q\|p)$. For populations, value is frame-relative while price is frame-independent; a fleet that pools its resource and fuses its perception inherits the ceiling $G_{\rm fleet}\le I(X;Y_{1:m})\le H(X)$ (a corollary; an earlier sum-form claim was wrong and is corrected in v5). A dynamical layer yields an is/ought asymmetry from which alignment emerges as a control-stability condition. We test the single-frame laws on live language models, pre-registered: perception mutual information tracks realized capability (Spearman $ρ=0.977$ over 30 model$\times$domain points); out-of-sample $ΔG$ tracks $I(X;Y)$, shape-invariant across four task shapes ($n=42$, slope $0.953$); over-confidence is measurable dissipation. The stated continuation gate has since been run (pre-registered, frontier-model population): the coupled capacity-region prediction -- growth-gap law, coalition submodularity with an XOR synergy control, joint ceiling, Kelly selection -- is confirmed within its frozen bands on real agents; the mean-field residual law $\|Vg\|/γ$ found no domain (populations hold no goal dispersion) and is retired to its mathematical scope. The contribution is the unification and the governance mapping that follows.

价值理论智能体信息论对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。