智能分配人工审核资源,让有限人力聚焦最不可靠的AI预测任务。
Allocating Human Oversight in AI-Enabled Analytics
- 基于置信度上界动态分配人工审核,边做边学任务可靠性。
- 在68个任务的真实调查中,将误差率从10%-12%降至2%-6%。
- 适合需要优化人力成本的AI决策系统,尤其当预测可靠性不均时。
企业日益将AI作为低成本预测层应用于客户决策流程,如需求感知、服务质量监控、产品测试和市场调研,但不同任务、产品与客户群体中AI信号的可靠性差异显著。因此仍需稀缺的人工验证(标签、审计、问卷或跟进测量)来锚定AI输出与真实情况。由于人工标注本身存在噪声,且在不同标注者间、甚至重复判断中波动,企业必须对每项任务收集并平均多个标注,导致成本高昂。本文研究在部署前未知可靠性的情况下,如何在众多AI辅助任务间分配有限的人工验证预算。我们将其建模为调优后的预测驱动推断问题:每个人工标注不仅提升AI估计精度,还揭示任务的修正难度(即最优使用AI作为协变量后仍残留的方差)。若难度已知,最优分配遵循奈曼平方根规则;因未知,我们提出一种基于上置信界的学习策略,实时学习难度并引导验证资源向AI最不可靠的任务倾斜。理论证明,该策略的最终效率损失随预算增加趋近于零。在合成实验和包含68个任务、超过2000名受访者的真实数字孪生调查中,该方法在可靠性异质时几乎逼近理想分配性能,优于均匀分配与ε-贪婪策略;在真实数据上亦超越探索-再利用试点设计,将均匀分配的10%-12%差距缩小至2%-6%。由此可见,AI的价值不仅取决于模型精度,更取决于能否将人工监督精准投向最需干预的环节。
原文摘要 · Abstract (English)
Organizations increasingly deploy AI as a low-cost prediction layer in customer-facing decision processes, including demand sensing, service-quality monitoring, product testing, and market research, but AI-generated signals are unevenly reliable across tasks, products, and customer segments. Firms therefore still need scarce human validation (labels, audits, survey responses, or follow-up measurements) to anchor AI outputs to ground truth. Because human ground truth is itself noisy, varying across labelers and even across repeated judgments, the firm must collect and average several human labels per task, which makes human validation costly. We study how to allocate a limited human-validation budget across many AI-assisted tasks when reliability is heterogeneous and unknown before deployment. We cast this within tuned prediction-powered inference. Each human label both sharpens the AI-assisted estimate and reveals the task's rectification difficulty, the variance that remains after the AI prediction is optimally used as a control variate. If difficulties were known, the optimal allocation would follow a Neyman square-root rule; because they are unknown, we propose a policy based on upper confidence bounds that learns them online and steers validation toward tasks where AI is least reliable. We prove that the policy's terminal efficiency loss relative to the oracle allocation vanishes as the budget grows. In synthetic experiments and a real digital-twin survey with 68 tasks and over 2000 respondents, it closes most of the gap to the oracle when reliability is heterogeneous, outperforming uniform and epsilon-greedy allocation; on the survey data it also outperforms explore-then-commit pilot designs and cuts uniform's 10--12% gap to 2--6%. The value of AI depends not only on model accuracy but also on the operational policy that targets human oversight where AI errors matter most.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。