arXiv:2605.00844cs.CYcs.AI2026-05被引 3

三大独立大模型预测错误高度相似,揭示潜在认知同质化风险。

The Oracle's Fingerprint: Correlated AI Forecasting Errors and the Limits of Bias Transmission

  • 三款独立大模型在568个预测题中错误相关性达0.77,显示共同偏差模式。
  • 人类群体预测受大模型影响,但仅因理性修正真实结果,未被模型偏见污染。
  • 人类原有偏见已与大模型模式高度一致,说明模型并非引入新偏见。

当大型语言模型(LLMs)被用作预测工具时,个体错误的独立性——集体智慧的基础——可能崩溃。我们测试了三种促成‘认知单一化’的条件。研究1显示,GPT-4o、Claude和Gemini在568个已解决的二元预测问题上表现出高度相关的预测错误(平均成对错误相关系数 r = 0.77,p < 0.001;排除可能泄露问题后 r = 0.78),尽管它们由不同机构独立开发。研究2通过跨问题设计,追踪了ChatGPT发布边界(2022年11月)前后社区预测的变化,发现群体预测朝向模型预测方向移动(r = 0.20,p = 0.007),但这种变化完全可由向真实答案的理性更新解释。研究3考察了人类预测错误的类别模式是否逐渐趋近于大模型的偏见指纹,结果相反:发布前的人类偏见已与大模型模式高度相似(r = 0.87),而发布后相似度下降至 r = -0.28。综合来看,一种已建立但尚未激活的认知单一化正在形成:三款名义上独立的AI系统共享相同失败模式,放大了人类固有的偏见。

原文摘要 · Abstract (English)

When large language models (LLMs) are consulted as forecasting tools, the independence of individual errors -- the foundation of collective intelligence -- may collapse. We test three conditions necessary for this "epistemic monoculture" to emerge. In Study 1, we show that GPT-4o, Claude, and Gemini exhibit highly correlated forecasting errors on 568 resolved binary prediction questions (mean pairwise error correlation r = 0.77, p < 0.001; r = 0.78 excluding likely-leaked questions), despite being developed independently by different organizations. In Study 2, we test whether this correlated bias has propagated into human crowd forecasts, using a within-question design that tracks community prediction shifts across the ChatGPT launch boundary (November 2022). We find that community forecasts move in the direction predicted by LLMs (r = 0.20, p = 0.007), but this shift is fully explained by rational updating toward ground truth. In Study 3, we examine whether the category-level pattern of human forecasting errors increasingly resembles the LLM bias fingerprint. We find the opposite: pre-ChatGPT human biases already strongly resembled the LLM pattern (r = 0.87), while post-ChatGPT the resemblance weakened (r = -0.28). Together, these findings reveal an epistemic monoculture that is built but not yet activated: three nominally independent AI systems share the same failure modes, amplifying precisely the biases humans already hold.

大模型认知偏见预测误差群体智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。