检测大模型群体分歧是否带来真实认知修正,避免表面多样实则固执。
When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Diagnostic for Machine Collectives
- 通过干预输出分散度,检验群体是否真改变观点而非仅换说法。
- GPT-4o-mini在分歧下假前提修复提升17.7分,Gemini却无改善。
- 发现Gemini用内部辩论伪装分歧,94%回应只是重述不认错。
集体智能研究将分歧视为认知多样性证据:若个体表达不同观点,群体应具备修正能力。但在大语言模型群体中,这种代理可能失效——模型可生成看似多样的论证,却始终维持相同结论。本文提出黑箱诊断方法,衡量输出分散性与认知立场真实修正之间的耦合程度:即干预是否真正引发认知改变,而非仅保留前提的重新表述。该诊断仅基于生成文本,不涉及模型内部表示。通过两个独立通道评估:输出通道使用一致性指数(CI)验证分散性变化;认知通道通过逐轮立场标注判断是否发生立场修正。提出基于元预测清晰度系统(MPCS)的再分化协议(RDP),用于动态提升输出分散性。在两种配置(gpt-4o-mini 和 gemini-2.5-flash)的五人集体中测试,每组310对实验。结果显示,gpt-4o-mini在条件分歧下假前提修复提升+17.7分(p<1e-6),而静态人格多样性反而降低恢复力(-8.1,p=.007)。gemini-2.5-flash在相近预算下未见提升(26.1% vs 27.1%,p=.84),尽管输出分散性下降;两组处理效果差异显著(z=3.79,p<.001)。机制标记显示,Gemini通过框架内分歧保持错误前提:94%的后协议回应为重述而非让步(对比GPT为24%)。建议报告每干预后的立场转变率和前提保留率。
原文摘要 · Abstract (English)
Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion. We operationalize dispersion-revision coupling: the degree to which an intervention that verifiably increases the dispersion of a collective's outputs in embedding space is accompanied by genuine revision of its epistemic stance rather than premise-preserving reformulation. The diagnostic is black-box: it operates on generated text alone and makes no claims about the internal representations of the generating models. Two channels are measured independently: an output channel, the Coherence Index (CI), verifies that the intervention changed output dispersion; an epistemic channel, per-turn stance annotation, measures whether the collective revised. We propose CI with the Meta-Predictive Clarity System (MPCS), which inserts a Re-Differentiation Protocol (RDP) when outputs over-converge, as a reusable method for estimating this coupling regime. We evaluate five-agent collectives from two configurations (gpt-4o-mini and gemini-2.5-flash; 310 paired episodes per condition). On gpt-4o-mini, conditional dissent improves false-premise recovery by +17.7 points (p<1e-6) while static persona diversity harms recovery (-8.1, p=.007). On gemini-2.5-flash, the same intervention at a comparable budget yields no gain (26.1% vs 27.1%, p=.84) despite a verified dispersion drop; the two treatment effects differ from each other (z=3.79, p<.001). Mechanism tagging shows Gemini preserves the false premise via intra-framework dissent: 94% of tagged post-RDP responses reformulate rather than concede (vs 24% on GPT). We recommend reporting per-intervention stance shift and premise-preservation rate alongside accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。