用错损失函数会误导AI,反而让医生更难做出正确判断。
Misaligned by Design: Incentive Failures in Machine Learning
- 用人类目标设计损失函数,反而削弱AI学习能力
- 事后调整预测比训练时对齐人类目标更有效
- 适合医疗决策等高风险场景的AI设计参考
在许多高风险场景中,错误成本具有不对称性:误诊肺炎虽不便,但漏诊可能致命。因此,辅助决策的AI模型常采用包含人类权衡取舍的非对称损失函数进行训练。然而,在两个关键应用中,我们发现这种标准对齐做法反而适得其反。更好的方法是训练模型时忽略人类目标,再事后根据人类目标调整预测。我们的经济激励模型表明,机器分类器需完成双重任务:如何分类和如何学习分类。现有调整机制虽能正确激励分类决策,却可能削弱学习动力。理论分析揭示,看似直观的方法反而以可预测方式造成人机目标错配。
原文摘要 · Abstract (English)
The cost of error in many high-stakes settings is asymmetric: misdiagnosing pneumonia when absent is an inconvenience, but failing to detect it when present can be life-threatening. Because of this, artificial intelligence (AI) models used to assist such decisions are frequently trained with asymmetric loss functions that incorporate human decision-makers' trade-offs between false positives and false negatives. In two focal applications, we show that this standard alignment practice can backfire. In both cases, it would be better to train the machine learning model with a loss function that ignores the human's objective and then adjust predictions ex post according to that objective. We rationalize this result using an economic model of incentive design with endogenous information acquisition. The key insight from our theoretical framework is that machine classifiers perform not one but two incentivized tasks: choosing how to classify and learning how to classify. We show that while the adjustments engineers use correctly incentivize choosing, they can simultaneously reduce the incentives to learn. Our formal treatment of the problem reveals that methods embraced for their intuitive appeal can in fact misalign human and machine objectives in predictable ways.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。