多智能体推理中,报告相似不等于证据独立,需识别来源依赖性。
Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence
- 提出'认知塞比问题',量化报告间真实证据关联性
- 实验显示报告量增至32倍时后验覆盖从0.940降至0.263
- 建议追踪证据源头而非报告数量或相似度,适合可靠性研究者
多智能体系统通过生成多个代理并整合报告来提升推理能力。但一个代理并不等同于一次独立观测:看似独立的报告可能源自同一证据源,而真正独立的证据也可能产生高度相似的报告。本文将此现象形式化为‘认知塞比问题’——当给定已有报告R时,新报告Z与参数Theta的互信息为零,则称其为认知塞比扩展。仅基于报告的聚合器无法区分重复与独立佐证,因相同报告在不同祖先关系下可对应不同后验概率。高斯共享根模型表明,共同起源未必导致完全冗余;重复提取可逼近源级上限,而共享基础模型引发的提取误差相关性会进一步压低该上限。我们在合成证据文档上进行了超过2万次受控的LLM代理报告与提取调用测试。固定一个证据源,报告数量从1增至32时,朴素后验覆盖率从0.940降至0.263;固定报告数,证据源数量从1增至16时,聚合器性能趋同,在k=16时统计无差异。代理间的重复提取误差高度相关(样本外估计γ=0.719),采用相关提取聚合器可恢复校准。受控操纵实验分离表示相似性与证据来源:报告空间去重机制的平均聚类数变化1.425(95% CI [1.363, 1.485]),而真实来源变化四倍仅引起0.040的变化([-0.045, 0.120])。因此,集体推理应关注证据来源与依赖结构,而非代理数量、报告数量或相似度。
原文摘要 · Abstract (English)
Multi-agent AI systems improve inference by spawning agents and synthesizing reports. But another agent is not another observation: apparently independent reports may descend from the same evidence, and genuinely independent evidence can produce nearly identical reports. We formalize this as an epistemic Sybil problem. A report Z is an epistemic Sybil extension relative to reports R when I(Theta; Z | R) = 0. No report-only aggregator can generally distinguish replication from independent corroboration: identical reports can warrant different posteriors under unobserved ancestry. A Gaussian shared-root model shows common ancestry does not imply complete redundancy. Repeated extraction adds information toward a source-level ceiling, and correlated extraction errors, which a shared base model can induce among independent agents, lower that ceiling further. We test these predictions with more than 20,000 controlled LLM-agent report and extraction calls on synthetic evidentiary documents. Holding one evidence root fixed while report multiplicity rises from 1 to 32 collapses naive posterior coverage from 0.940 to 0.263. Holding report count fixed while evidence-root multiplicity rises from 1 to 16 closes the gap, and the aggregators are statistically indistinguishable at k = 16. The agent's replicate extraction errors are correlated (gamma_cal = 0.719, estimated out of sample), and a correlated-extraction aggregator restores calibration accordingly. A controlled manipulation isolates representation similarity from evidential ancestry. It changes a report-space deduplication mechanism's mean inferred cluster count by 1.425 (95% CI [1.363, 1.485]), whereas a fourfold change in true ancestry changes it by only 0.040 ([-0.045, 0.120]). Collective inference should therefore track evidential ancestry and dependence, not agent or report multiplicity or similarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。