用概率模型揭示大模型中潜在代理结构的共存与对齐机制
Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks
- 将智能体建模为带认知效用的输出分布,通过加权对数池化实现协作优化
- 三类以上结果空间下可实现严格一致,而二元空间或线性池化无法达成
- 实证发现大模型中善恶人格互斥共现,压制策略比单纯强化更有效
我们构建了一个基于概率建模的智能体理论,将智能体表示为具有认知效用(以对数得分衡量)的输出分布,其组合通过加权对数池化实现,且严格提升所有成员福利。证明在二元结果空间或线性池化下严格一致不可能实现,但在三类及以上结果空间下可能。框架通过克隆不变性、连续性和开性支持递归结构,倾斜分析排除平凡复制。最后,用该理论形式化了大语言模型中的代理对齐现象:诱导仁慈人格('Luigi')会激发对抗人格('Waluigi'),而先显现再压制'Waluigi'的策略,相较单纯强化'Luigi',能实现更显著的一阶对齐提升。这些结果表明,建立子代理如何聚合为高层实体的数学框架,为智能体系统的对齐提供了新洞见。
原文摘要 · Abstract (English)
We develop a theory of intelligent agency grounded in probabilistic modeling for neural models. Agents are represented as outcome distributions with epistemic utility given by log score, and compositions are defined through weighted logarithmic pooling that strictly improves every member's welfare. We prove that strict unanimity is impossible under linear pooling or in binary outcome spaces, but possible with three or more outcomes. Our framework admits recursive structure via cloning invariance, continuity, and openness, while tilt-based analysis rules out trivial duplication. Finally, we formalize an agentic alignment phenomenon in LLMs using our theory: eliciting a benevolent persona ("Luigi'") induces an antagonistic counterpart ("Waluigi"), while a manifest-then-suppress Waluigi strategy yields strictly larger first-order misalignment reduction than pure Luigi reinforcement alone. These results clarify how developing a principled mathematical framework for how subagents can coalesce into coherent higher-level entities provides novel implications for alignment in agentic AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。