校准模型预测后,人类决策仍不靠谱,需结合行为经济学修正才有效。
Does Calibration Affect Human Actions?
- 用行为经济学修正校准结果,提升人类决策与模型预测的一致性。
- 修正后人类决策与预测相关性显著提高,但信任感未明显变化。
- 非专家用户依赖模型时,仅校准不够,需心理机制干预。
校准被提出用于提升机器学习分类器的可靠性与可接受度。我们研究了这一提议的一个特定方面:校准分类模型如何影响非专家用户对模型预测的决策行为。通过人机交互实验,考察了校准对(i)用户对模型的信任度、(ii)决策与预测间相关性的影响。我们还基于卡尼曼与特沃斯基的行为经济学前景理论,提出对校准分数的进一步修正,并研究其对信任与决策的影响。结果表明,仅校准本身不足以改善效果;前景理论修正对于提升人类决策与模型预测的相关性至关重要。尽管这种相关性增强暗示更高信任,但用户在‘你是否更信任模型?’问题上的回答并未因方法不同而改变。
原文摘要 · Abstract (English)
Calibration has been proposed as a way to enhance the reliability and adoption of machine learning classifiers. We study a particular aspect of this proposal: how does calibrating a classification model affect the decisions made by non-expert humans consuming the model's predictions? We perform a Human-Computer-Interaction (HCI) experiment to ascertain the effect of calibration on (i) trust in the model, and (ii) the correlation between decisions and predictions. We also propose further corrections to the reported calibrated scores based on Kahneman and Tversky's prospect theory from behavioral economics, and study the effect of these corrections on trust and decision-making. We find that calibration is not sufficient on its own; the prospect theory correction is crucial for increasing the correlation between human decisions and the model's predictions. While this increased correlation suggests higher trust in the model, responses to ``Do you trust the model more?" are unaffected by the method used.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。