为人类与智能体协作流程设计最优评估方案,提升决策准确性。
Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era
- 提出双概率重播策略,动态分配人类或代理单独测试的优先级。
- 在临床和编码任务中验证:仅部分场景下协作流程优于单一模式。
- 适合需要权衡人力与算力成本的智能系统部署决策者使用。
当编码代理受工程师监督,或临床模型辅助放射科医生时,关键问题是:应保留人机协作流程,还是改用纯人类或纯代理?只有当协作流程优于两者时才值得保留。然而部署后无法观察到替代方案的表现,需通过重跑任务来补全数据,但每次重跑都消耗专家时间或计算资源。在固定重跑预算下,核心问题是:哪些任务应更优先进行纯人类重跑,哪些应进行纯代理重跑?现有方法未直接解决此决策问题:基准评测不选择缺失基线,方差采样忽略两组对比中哪一组更难判定,贝叶斯信息方法关注参数学习而非部署决策。本文提出TEAM-Design规则,为每项任务分配两个重跑概率,分别对应人类或代理基线。该规则在已知信息难以预测基线结果、且对比更难确立时提高重跑概率,在重跑成本高时降低概率。理论证明该规则能最优解决带预算限制的设计问题;随机按记录概率抽样仍能控制错误宣称协作流程胜出的概率。我们重新分析6个临床场景(均无协作流程胜过两者),以及一个编码基准(有一个胜出),并在合成设计与基于真实胸部X光读片研究的半合成设计上评估了该方法。结果显示:当两个对比难度差异明显时,该方法表现最佳;当两者难度相近时,可能不如方差采样。
原文摘要 · Abstract (English)
Wherever a coding agent works under engineer supervision, or a clinical model assists a radiologist, the deployment question is whether to keep the human-AI workflow or replace it with the human alone or the agent alone. The human-AI workflow is worth keeping only if it beats both of those alternatives. Yet once it is deployed, neither alternative outcome is observed: recovering one means replaying the task under that alternative, and every replay costs expert time or compute. Under a fixed replay budget, the design question is therefore which tasks should be more likely to receive a human-only replay, and which an agent-only replay. Existing methods do not directly target this decision. Agent benchmarks do not choose which missing baseline to measure, variance-based sampling ignores which of the two comparisons is closer to failing, and Bayesian information methods focus on learning model parameters instead of making the deployment decision. We propose TEAM-Design, a rule that gives every task two replay probabilities, one per baseline. It raises a probability where the missing baseline outcome is hard to predict from what is already known about the task and where that comparison is harder to establish, and lowers it where replay is expensive. We prove that the rule solves this budgeted design problem, and that drawing the replays at random from recorded probabilities still controls the chance of wrongly declaring that the workflow beats both. We reanalyze 6 clinical settings, where no human-AI workflow beats both alternatives, and a coding benchmark, where one does, then evaluate TEAM-Design on synthetic designs and on a semi-synthetic design built from a real chest X-ray reader study. TEAM-Design works best when one of the two comparisons is clearly harder to settle than the other, and can do worse than variance-based allocation when the two are similarly difficult.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。