arXiv:2603.06197cs.CL2026-03

用多个大模型集体判断,逼近内容分析的真相。

Wisdom of the AI Crowd (AI-CROWD) for Ground Truth Approximation in Content Analysis: A Research Protocol & Validation Using Eleven Large Language Models

  • 让11个大模型集体打标签,通过投票和分歧分析找共识。
  • 能识别高置信度结果,同时标记模糊或偏见区域。
  • 适合没人工标注、需快速建模的大规模内容分析。

大规模内容分析日益受限于缺乏可观察的真实标签(金标准),因人工标注海量数据耗时长、成本高且难以保证一致性。为突破此瓶颈,我们提出AI-CROWD协议,通过集成多个大语言模型(LLMs)的输出,近似生成真实标签。该协议不宣称结果为绝对真值,而是基于多模型间收敛与分歧推断,形成共识性近似。通过多数投票及诊断性指标分析一致/不一致模式,识别出高置信度分类,并标示潜在模糊性或模型特异性偏差,适用于缺乏人工标注的场景。

原文摘要 · Abstract (English)

Large-scale content analysis is increasingly limited by the absence of observable ground truth or gold-standard labels, as creating such benchmarks through extensive human coding becomes impractical for massive datasets due to high time, cost, and consistency challenges. To overcome this barrier, we introduce the AI-CROWD protocol, which approximates ground truth by leveraging the collective outputs of an ensemble of large language models (LLMs). Rather than asserting that the resulting labels are true ground truth, the protocol generates a consensus-based approximation derived from convergent and divergent inferences across multiple models. By aggregating outputs via majority voting and interrogating agreement/disagreement patterns with diagnostic metrics, AI-CROWD identifies high-confidence classifications while flagging potential ambiguity or model-specific biases.

内容分析大模型集成标签近似共识机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。