arXiv:2604.26561cs.MAcs.AI2026-04

让AI辩论更真实:用不同模型和一致性验证避免虚假共识

Preserving Disagreement: Architectural Heterogeneity and Coherence Validation in Multi-Agent Policy Simulation

  • 为不同价值立场分配不同参数的AI模型,减少观点趋同
  • 异构模型使首选方案集中度下降,如儿童福利从70.9%降至46.1%
  • 一致性验证有双面效果:在主导选项下减少集中,竞争选项下反而加剧

基于大语言模型的多智能体政策模拟系统常出现虚假共识问题:评估智能体无论立场如何都趋向同一选项。本文提出AI理事会三阶段框架,在两个政策场景中开展120次模拟,测试两种干预。首先,架构异质性(为每个价值立场分配不同7-9B参数模型)显著降低首选项集中度(儿童福利:70.9%→46.1%,p<0.001,r=0.58;住房:46.0%→22.9%,p<0.001,r=0.50)。此效果在无客观正确答案时存在,与追求准确性的辩论模式不同。其次,一致性验证(使用前沿模型检验推理是否符合设定立场)揭示了真实性与多样性权衡:在主导选项场景中进一步降低集中度(46.1%→40.8%,p=0.004),但在真正竞争性选项场景中反而提升集中度(22.9%→26.6%,p=0.96),因高一致性智能体聚集于单一选项。该权衡可能普遍存在于质量加权的多智能体系统中。报告三个失败的德尔菲设计结果,发现8B模型对反驳呈二元响应而非渐进式,提出可信张力率作为小模型辩论能力诊断指标。

原文摘要 · Abstract (English)

Multi-agent deliberation systems using large language models (LLMs) are increasingly proposed for policy simulation, yet they suffer from artificial consensus: evaluator agents converge on the same option regardless of their assigned value perspectives. We present the AI Council, a three-phase deliberation framework, and conduct 120 deliberations across two policy scenarios to test two interventions. First, architectural heterogeneity (assigning a different 7-9B parameter model to each value perspective) significantly reduces first-choice concentration compared to a homogeneous baseline (child welfare: 70.9% to 46.1%, p < 0.001, r = 0.58; housing: 46.0% to 22.9%, p < 0.001, r = 0.50). This contrasts with accuracy-oriented multi-agent debate, where heterogeneity does not reduce convergence, suggesting model diversity operates differently when no objectively correct answer exists. Second, coherence validation (using a frontier model to assess whether each evaluator's reasoning is grounded in its assigned values) reveals a fidelity-diversity tradeoff: on a scenario with a dominant option, it further reduces concentration (46.1% to 40.8%, p = 0.004), but on a scenario with genuinely competitive options, it increases concentration (22.9% to 26.6%, p = 0.96) by amplifying high-coherence evaluators who cluster on one option. This tradeoff may be a general property of multi-agent systems employing quality weighting. We report negative results from three failed Delphi designs, demonstrate that 8B models exhibit binary rather than graded responses to counter-arguments, and propose the trustworthy tension rate as a diagnostic measure of small-model deliberation capabilities.

多智能体政策模拟共识偏差模型异质

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。