研究人机协作中校准机制如何影响决策效果。
Human-AI Teaming Through the Lens of Calibration

- 基于特征空间划分,分析人与AI的校准如何影响协作
- 现有融合方法会破坏人类预测的校准性
- 委托机制需精细校准决策模型,对专家型人类挑战大
我们从统计校准的角度研究人机协作模型。假设团队由人类和AI组成,二者均在特征空间的某种划分下具备校准性,揭示校准假设如何传导至协作框架。具体考虑两种机制:(i) 融合人类与模型的预测结果;(ii) 将预测责任分配给其中一方。理论与实证结果表明,现有融合方法无法保持人类的校准程度;而委托机制虽保留下游预测者的校准性,却将判断责任转移至拒绝决策的元模型。该元模型必须足够精细地识别出双方优势区域,其要求随人类能力提升而增加,当人类依赖系统无法观测的信息时,此需求变得不可实现。
原文摘要 · Abstract (English)
We study models for human-AI teaming through the lens of statistical calibration. We assume the team consists of an AI model and human -- both of which are calibrated with respect to some partitioning of the feature space -- and expose how the calibration assumptions propagate into the teaming framework. In particular, we consider frameworks that either (i) combine human and model predictions or (ii) delegate prediction responsibility to either a human or model. We show via theoretical and empirical results that existing methods for combination do not preserve the human's degree of calibration. Methods for delegation (by the very act of delegation) preserve calibration of the downstream predictors but shift the burden onto the rejector meta-model that decides who predicts. The rejector must be calibrated finely enough to locate where each member is superior, a demand that grows with the human's expertise and becomes unattainable when the human relies on information the system cannot observe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。