arXiv:2606.31404cs.AI2026-06中稿 · ECIS 2026

用大模型模拟群体智慧,显著提升预测准确率。

Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models

  • 让多个大模型协作估测,通过聚合降低误差。
  • 最高误差减少37个百分点,效果随策略优化提升。
  • 模型能自知不确定,适合用于决策支持系统。

人类群体智慧虽具高准确性,但受限于成本、协调与时间。本文探究大语言模型(LLMs)能否通过人工群体模拟群体智能,填补了基于AI的聚合机制研究空白。在三个专有模型(GPT-5、Gemini 2.5 Pro、Claude Sonnet 4.5)上,通过960次手动提示,测试了模型内采样与跨模型聚合在八项估测任务中的表现。结果表明,模型内与跨模型聚合均带来一致的误差下降,不同策略下平均绝对百分比误差(MAPE)最高降低37个百分点。观察到相对置信区间宽度与相对估测误差之间存在小至中等程度的正相关(Spearman's ρ=0.242–0.568,所有p<0.001),表明大模型具备评估不确定性的元认知能力。研究为学术与实践提供可操作洞察,推动大模型群体在组织决策中的应用。

原文摘要 · Abstract (English)

Human swarm intelligence demonstrates remarkable collective accuracy but faces scalability constraints in cost, coordination, and time. We investigate whether large language models (LLMs) can approximate swarm intelligence effects through artificial swarms, addressing a critical gap in understanding AI-based aggregation mechanisms. We conducted a controlled experiment with 960 manually executed prompts across three proprietary models (GPT-5, Gemini 2.5 Pro, Claude Sonnet 4.5), testing intra-model sampling and inter-model aggregation on eight estimation tasks. Results reveal consistent error reduction through intra- and inter-model aggregation, with significant error reductions up to 37 percentage points in MAPE across different aggregation strategies. We observed small to large effect sizes for positive correlations (Spearman's $ρ=0.242-0.568$, all $p<0.001$) between relative confidence interval widths and relative estimation errors, suggesting LLMs possess metacognitive awareness when assessing uncertainty. We discuss implications for research and practice, providing actionable insights for deploying LLM swarms in organizational decision-making.

群体智能大模型决策支持元认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。