arXiv:2602.14471cs.MAcs.AI2026-02被引 2

让AI代理学会权衡个人利益与集体稳定,避免系统拥堵

Socially-Weighted Alignment: A Game-Theoretic Framework for Multi-Agent LLM Systems

  • 用社会权重动态调节代理的决策,平衡私利与群体福祉
  • 当社会权重超过阈值λ*=(n−β)/(n−1)时,系统从拥堵转为稳定
  • 无需训练即可应用,适合多智能体协作场景

在共享环境中部署大语言模型代理时,个体理性决策可能引发负面外部性,损害整体性能。我们提出社交加权对齐(SWA),一种基于博弈论的推理期框架,通过社会权重λ∈[0,1]在代理私有目标与群体福利估计间插值来调整决策。在包含n个代理、拥堵严重度为β的共享资源拥堵游戏中,我们证明:当λ超过临界阈值λ*=(n−β)/(n−1)时,代理在过载情况下不再有增加需求的边际动机,系统发生相变,从持续拥堵转向接近容量的稳定运行。我们进一步设计了无需参数更新或多智能体强化学习的推理期算法实现,并通过多智能体模拟验证了预测的阈值行为。

原文摘要 · Abstract (English)

Deploying large language model (LLM) agents in shared environments introduces a fundamental tension between individual alignment and collective stability: locally rational decisions can impose negative externalities that degrade system-level performance. We propose Socially-Weighted Alignment (SWA), a game-theoretic framework that modifies inference-time decision making by interpolating between an agent's private objective and an estimate of group welfare via a social weight $λ\in[0,1]$. In a shared-resource congestion game with $n$ agents and congestion severity $β$, we show that SWA induces a critical threshold $λ^*=(n-β)/(n-1)$ above which agents no longer have marginal incentive to increase demand under overload, yielding a phase transition from persistent congestion to stable operation near capacity. We further provide an inference-time algorithmic instantiation of SWA that does not require parameter updates or multi-agent reinforcement learning, and use a multi-agent simulation to empirically validate the predicted threshold behavior.

多智能体博弈论大模型对齐系统稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。