测试人类与AI协作在多种任务中的互补性,发现效果有限且依赖精准决策路由。
Toward Human-AI Complementarity Across Diverse Tasks

- 用混合判断和两种AI辅助方法测试跨任务互补性
- 人类仅提升0.4个百分点,低置信度时辅助使准确率从28.4%升至38.3%
- 关键瓶颈是无法有效识别该由人介入的错误场景
人类与AI的互补性——即两者结合判断优于单独使用——为先进AI系统的可靠监管提供了潜在路径。但这一能力在真实任务中是否成立仍不确定。我们通过混合判断及两种AI辅助方法(前两名建议、子任务委派),在涵盖知识、事实性、长上下文推理和欺骗检测的1,886样本多领域数据集上进行评估。结果表明互补性提升有限:基线混合仅比纯AI高0.4个百分点(69.3% vs 68.9%),且互补区域极小(仅8.9%的样本中AI出错而人类正确)。信心路由无效,因模型在正确与错误预测上的置信度分布相似。当AI置信度低时,前两名建议使人类准确率从28.4%提升至38.3%,超过AI单独表现(37.7%),但主要得益于采纳正确建议,而非纠正错误。分析揭示核心瓶颈并非人类能力不足,而是决策路由精度与辅助设计能否有效捕捉AI失误。研究为未来合作优化提供明确方向,数据与代码将按需公开。
原文摘要 · Abstract (English)
Human-AI complementarity, the idea that combining human and AI judgments can outperform either alone, offers a promising pathway toward robust oversight of advanced AI systems. However, whether human-AI complementarity can be achieved on realistic tasks remains an open question. We investigate this through two approaches: hybridization and two AI assistance methods (top-2 assistance and subtask delegation), evaluated on a multi-domain dataset of 1,886 samples spanning knowledge, factuality, long-context reasoning, and deception detection. We find only modest complementarity gains. Baseline hybridization yields just +0.4 percentage points (pp) over AI alone (69.3\% vs 68.9\%), limited both by a small complementarity region (only 8.9\% of items where AI errs but humans do not) and the inability of confidence-based routing to identify it, since the model's confidence is similarly distributed across correct and incorrect predictions. Applied when AI has low confidence, top-2 assistance increases human accuracy from 28.4\% to 38.3\%, surpassing AI alone (37.7\%) -- but primarily because humans adopt correct AI suggestions, not because they successfully override AI errors. These findings suggest that the primary bottleneck is not human task accuracy per se, but the ability to route decisions to humans when it matters and to design assistance methods that enable humans to catch AI mistakes. Our quantitative and qualitative analyses pinpoint where and why each method succeeds or fails, offering concrete targets for future work. We will release our dataset and code upon request to support progress toward more effective human-AI collaboration for AI oversight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。