arXiv:2510.23476cs.AIcs.HC2025-10被引 8

让人类与AI协作量化不确定性,提升决策可靠性。

Human-AI Collaborative Uncertainty Quantification

  • 设计协同框架,用AI优化人类预测集,避免错误且补全遗漏。
  • 实验显示协作结果覆盖更高、集合更小,优于单独使用人或AI。
  • 算法可应对数据分布变化,适合医疗等高风险决策场景。

AI预测系统正广泛嵌入高风险决策流程,但面对不确定性时,当前AI仍缺乏人类具备的领域知识、长期上下文理解及物理世界推理能力。为此,本文提出人类-人工智能协同不确定性量化框架(Human-AI Collaborative Uncertainty Quantification),旨在通过AI对人类提出的预测集进行优化:一、避免反事实伤害,不削弱正确的人类判断;二、实现互补性,恢复人类遗漏的正确结果。在群体层面,我们证明最优协作预测集遵循单一得分函数上的双阈值结构,扩展了经典共形预测理论。基于此,开发了具有无分布有限样本保证的离线与在线校准算法。在线方法能适应分布漂移,包括人类在与AI交互中行为演化(称作“人向AI适应”)。在图像分类、回归及基于文本的医疗决策任务中,协作预测集始终优于单方表现,覆盖更高、集合更小,适用于多种条件。

原文摘要 · Abstract (English)

AI predictive systems are increasingly embedded in decision making pipelines, shaping high stakes choices once made solely by humans. Yet robust decisions under uncertainty still rely on capabilities that current AI lacks: domain knowledge not captured by data, long horizon context, and reasoning grounded in the physical world. This gap has motivated growing efforts to design collaborative frameworks that combine the complementary strengths of humans and AI. This work advances this vision by identifying the fundamental principles of Human AI collaboration within uncertainty quantification, a key component of reliable decision making. We introduce Human AI Collaborative Uncertainty Quantification, a framework that formalizes how an AI model can refine a human expert's proposed prediction set with two goals: avoiding counterfactual harm, ensuring the AI does not degrade correct human judgments, and complementarity, enabling recovery of correct outcomes the human missed. At the population level, we show that the optimal collaborative prediction set follows an intuitive two threshold structure over a single score function, extending a classical result in conformal prediction. Building on this insight, we develop practical offline and online calibration algorithms with provable distribution free finite sample guarantees. The online method adapts to distribution shifts, including human behavior evolving through interaction with AI, a phenomenon we call Human to AI Adaptation. Experiments across image classification, regression, and text based medical decision making show that collaborative prediction sets consistently outperform either agent alone, achieving higher coverage and smaller set sizes across various conditions.

不确定性量化人机协作医疗决策共形预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。