arXiv:2605.11789cs.AI2026-05

用大模型模拟辩论,发现不文明沟通让讨论效率下降25%。

Beyond Inefficiency: Systemic Costs of Incivility in Multi-Agent Monte Carlo Simulations

论文配图:Beyond Inefficiency: Systemic Costs of Incivility in Multi-Agent Monte Carlo Simulations
图 1 · 摘自论文原文
  • 用大模型构建可控辩论环境,系统测试不同冒犯程度对效率影响。
  • 不文明沟通使达成结论的轮次增加25%,小模型受影响更严重。
  • 先发言者有显著优势,无论是否激烈争论,胜率均远超随机。

不建设性的争论和不文明沟通会显著损害效率与协作,但其对操作效率的具体影响难以量化。由于伦理限制、可复现性差及自然情境的不可预测性,人类实验难以深入研究。本文利用基于大语言模型(LLM)的多智能体系统,构建受控的社会学实验环境,大规模操控沟通行为。通过蒙特卡洛模拟框架,生成数千组一对一对抗性辩论,在不同毒性条件下测量收敛时间(达成结论所需轮次),作为互动效率的代理指标。在延续先前研究的基础上,扩展至两个参数量不同的新LLM智能体,验证结果是否具有跨模型规模的普适性。结果显示,此前报告的25%收敛延迟得到确认,且参数较少的模型受负面影响更大。此外,还发现明显的先发优势:发起对话的智能体获胜率显著高于随机水平,无论沟通是否具攻击性。

原文摘要 · Abstract (English)

Unconstructive debate and uncivil communication carry well-documented costs for productivity and cohesion, yet isolating their effect on operational efficiency has proven difficult. Human subject research in this domain is constrained by ethical oversight, limited reproducibility, and the inherent unpredictability of naturalistic settings. We address this gap by leveraging Large Language Model (LLM) based Multi-Agent Systems as a controlled sociological sandbox, enabling systematic manipulation of communicative behavior at scale. Using a Monte Carlo simulation framework, we generate thousands of structured 1-on-1 adversarial debates across varying toxicity conditions, measuring convergence time, defined as the number of rounds required to reach a conclusion, as a proxy for interactional efficiency. Building on a prior study, we replicate and extend its findings across two additional LLM agents of varying parameter size, allowing us to assess whether the effects of toxic behavior on debate dynamics generalize across model scale. The convergence latency of 25% reported in the previous study was confirmed. It was found that this latency is significantly bigger for models with fewer parameters. We further identify a significant first-mover advantage, whereby the agent initiating the discussion wins significantly above chance regardless of toxicity condition.

多智能体大模型对话效率社会模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。