用博弈论预测并引导大模型群体行为,避免政治偏见。
LLM Active Alignment: A Nash Equilibrium Perspective
- 将模型行为建模为对人类子群体的主动选择,形成可解释策略。
- 在社交媒体场景中,可防止模型群体排斥特定人群。
- 适合研究多模型协同与社会公平的开发者使用。
我们提出一种基于博弈论的框架,通过纳什均衡分析预测并引导大规模语言模型(LLMs)群体的行为。为避免在开放文本空间中均衡计算的不可行性,我们将每个智能体的动作建模为对人类子群体的混合选择。智能体主动且策略性地选择对齐对象,从而生成可解释且具有行为意义的策略类。在标准凹效用假设下,我们推导出闭式纳什均衡表征,实现系统级的解析预测,并提供明确、可操作的指导,帮助将对齐目标转向社会理想结果。该方法可作为现有对齐流程(如RLHF)之上的主动对齐层。在社交媒体场景中,我们发现,尤其是推理型模型组成的群体可能表现出政治排斥现象,即某些子群体被所有模型忽略;而本方法可有效避免此类问题,展示了其在跨领域调节多智能体大模型动态中的潜力。
原文摘要 · Abstract (English)
We develop a game-theoretic framework for predicting and steering the behavior of populations of large language models (LLMs) through Nash equilibrium (NE) analysis. To avoid the intractability of equilibrium computation in open-ended text spaces, we model each agent's action as a mixture over human subpopulations. Agents choose actively and strategically which groups to align with, yielding an interpretable and behaviorally substantive policy class. We derive closed-form NE characterizations, adopting standard concave-utility assumptions to enable analytical system-level predictions and give explicit, actionable guidance for shifting alignment targets toward socially desirable outcomes. The method functions as an active alignment layer on top of existing alignment pipelines such as RLHF. In a social-media setting, we show that a population of LLMs, especially reasoning-based models, may exhibit political exclusion, pathologies where some subpopulations are ignored by all LLM agents, which can be avoided by our method, illustrating the promise of applying the method to regulate multi-agent LLM dynamics across domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。