arXiv:2602.22413cs.AI2026-02

让模型自判可信度,不靠谱就别投票,提升集体决策准确率。

Epistemic Filtering and Collective Hallucination: A Jury Theorem for Confidence-Calibrated Agents

  • 模型通过自我校准学习可信度,低信心时选择不参与投票。
  • 理论证明:群体正确率在有限样本下仍能保证高于阈值。
  • 适合用于防范大模型集体幻觉,提升AI决策安全性。

我们研究了异质代理人在动态学习自身可靠性并选择性弃权情况下的集体准确性。传统投票理论如康多塞陪审团定理(CJT)假设固定参与,而现实中的聚合常受益于允许代理说“我不知道”。本文提出一个概率框架:代理先经历校准阶段,更新对自己固有能力的信念,再通过最终的置信度阈值决定是否投票或弃权。我们推导出群体成功概率的非渐近下界,并证明这种选择性参与将CJT的渐近保证推广到序列化、置信度门控的设定。通过蒙特卡洛模拟验证了这些边界。尽管结果具有普遍性,我们进一步讨论其在人工智能安全中的应用,说明该框架如何缓解大语言模型集体决策中的幻觉问题。

原文摘要 · Abstract (English)

We investigate the collective accuracy of heterogeneous agents who learn to estimate their own reliability over time and selectively abstain from voting. While classical epistemic voting results, such as the \textit{Condorcet Jury Theorem} (CJT), assume fixed participation, real-world aggregation often benefits from allowing agents to say ``I don't know.'' We propose a probabilistic framework where agents engage in a \textit{calibration} phase, updating beliefs about their own fixed competence, before facing a final confidence gate that determines whether to vote or abstain. We derive a non-asymptotic lower bound on the group's success probability and prove that this \textit{selective participation} generalizes the asymptotic guarantees of the CJT to a sequential, confidence-gated setting. Empirically, we validate these bounds via Monte Carlo simulations. While our results are general, we discuss their potential application to AI safety, outlining how this framework can mitigate \textit{hallucinations} in collective LLM decision-making.

群体决策模型校准幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。