用多智能体辩论模拟社会实验,量化AI的协作与认知行为。
The Social Laboratory: A Psychometric Framework for Multi-Agent LLM Evaluation
- 以角色化智能体辩论为社会实验室,观察其互动行为。
- 智能体自发达成高共识(μ > 0.88),即使在敏感话题上。
- 角色设定和主持者影响显著,适合评估对齐与社会行为。
随着大语言模型从静态工具演变为自主智能体,传统下游任务评估已难以捕捉其在交互环境中产生的社会与认知动态。为此,我们提出一种新评估框架,将多智能体辩论作为受控的‘社会实验室’,以发现并量化这些行为。在该框架中,具有不同人格和激励的LLM智能体,在LLM主持人监督下就广泛难题展开讨论。通过一套新的心理测量与语义指标分析发现:在数百场辩论中,智能体展现出强烈的自发共识倾向,达到高水平语义一致性(μ > 0.88),且无需明确指令;角色设定可诱发稳定可测的心理特征,尤其体现在认知努力上;主持人角色能显著改变辩论结果,这对外部AI对齐具有关键意义。本研究为下一代智能体提供了动态、基于心理测量的评估方法论。代码与结果已开源。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) transition from static tools to autonomous agents, traditional evaluation benchmarks that measure performance on downstream tasks are becoming insufficient. These methods fail to capture the emergent social and cognitive dynamics that arise when agents communicate, persuade, and collaborate in interactive environments. To address this gap, we introduce a novel evaluation framework that uses multi-agent debate as a controlled "social laboratory" to discover and quantify these behaviors. In our framework, LLM-based agents, instantiated with distinct personas and incentives, deliberate on a wide range of challenging topics under the supervision of an LLM moderator. Our analysis, enabled by a new suite of psychometric and semantic metrics, reveals several key findings. Across hundreds of debates, we uncover a powerful and robust emergent tendency for agents to seek consensus, consistently reaching high semantic agreement (μ > 0.88) even without explicit instruction and across sensitive topics. We show that assigned personas induce stable, measurable psychometric profiles, particularly in cognitive effort, and that the moderators persona can significantly alter debate outcomes by structuring the environment, a key finding for external AI alignment. This work provides a blueprint for a new class of dynamic, psychometrically grounded evaluation protocols designed for the agentic setting, offering a crucial methodology for understanding and shaping the social behaviors of the next generation of AI agents. We have released the code and results at https://github.com/znreza/multi-agent-LLM-eval-for-debate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。