揭示人机团队超越个体的条件,给出可验证的理论边界。
When Can Human-AI Teams Outperform Individuals? Tight Bounds with Impossibility Guarantees
- 基于置信度聚合规则,结合信号检测与信息论推导出互补性条件。
- 团队优势仅当错误相关性低于阈值时成立,且增益与元认知差异呈平方根关系。
- 提出不可行性定理,适用于非交互式决策,指导实际系统设计。
人机团队在70%的研究中未能超越其最优成员,但现有理论无法明确互补性的实现条件。本文通过整合信号检测理论与信息论分析,针对广泛的置信度聚合规则推导出紧致边界,得到四项结果:(1) 互补性定理——团队优于个体当且仅当错误相关性 $ρ_{HM} < ρ^*$,其中 $ρ^* o a$ 在对称近随机情形下;(2) 极小极大界表明收益规模为 $Θ( ext{√Δd})$,取决于元认知敏感性差异;(3) 不可行性结果证明当 $ρ_{HM} ≥ ρ^*$ 时,任何置信度聚合规则均无法实现互补性;(4) 多分类推广得 $ρ^*_K ≈ ρ^*/√{K−1}$。预测与实证高度吻合:ImageNet-16H 上 $R = 0.94$,CIFAR-10H 上 $R = 0.91$,多分类阈值缩放在人类数据中亦成立($R = 0.93$, $K = 16$),且对非高斯分布具有鲁棒性。该框架解释了互补性为何罕见,并提供可操作的设计公式,但仅适用于聚合决策,不适用于生成新答案的交互式讨论。
原文摘要 · Abstract (English)
Human-AI teams fail to outperform their best member in 70% of studies, yet no theory specifies when complementarity is achievable. We derive tight bounds for the broad class of confidence-based aggregation rules by integrating signal detection theory with information-theoretic analysis, yielding four results: (1) a complementarity theorem (teams outperform individuals iff error correlation $ρ_{HM} < ρ^*$, with $ρ^* \approx a$ in the symmetric near-chance regime); (2) minimax bounds showing gains scale as $Θ(\sqrt{Δd})$ with metacognitive sensitivity difference; (3) an impossibility result proving no confidence-based aggregation rule achieves complementarity when $ρ_{HM} \geq ρ^*$; and (4) multi-class generalization $ρ^*_K \approx ρ^*/\sqrt{K-1}$. Predictions match observed team accuracy ($R = 0.94$ on ImageNet-16H, $R = 0.91$ on CIFAR-10H) and the multi-class threshold scaling holds on human data ($R = 0.93$, $K = 16$), with robustness under non-Gaussian distributions. The framework explains why complementarity is rare and provides actionable design formulas; results apply to aggregation, not to interactive deliberation that generates novel answers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。