人能通过经验学会纠正对AI信心信号的误判。
Learning to Trust: How Humans Mentally Recalibrate AI Confidence Signals
- 用行为实验和计算模型研究人类如何调整对AI信心的信任。
- 50轮试验后,参与者信任判断准确率显著提升。
- 对反向信心信号适应困难,体现认知边界。
高效的人机协作需要恰当的信任,但当前AI系统常存在系统性过自信或不自信问题。我们研究了人类是否可通过重复经验来心理校准AI信心信号。在一项包含200名参与者的实验中,受试者需预测四种不同校准状态的AI(标准、过自信、不自信及反直觉的反向信心映射)的正确性。结果表明,在所有条件下,参与者经过50轮试验后,其预测准确性、区分度与校准一致性均显著提高。我们提出一个基于对数几率线性变换(LLO)与Rescorla-Wagner学习规则的计算模型,揭示人类通过更新基础信任水平和信心敏感度进行适应,且使用非对称学习率优先处理最具信息量的错误。尽管人类可补偿单调性偏差,但在反向信心情境下,大量参与者难以克服初始归纳偏见。该研究提供了人类如何通过经验调整对AI信心信任的机制解释。
原文摘要 · Abstract (English)
Productive human-AI collaboration requires appropriate reliance, yet contemporary AI systems are often miscalibrated, exhibiting systematic overconfidence or underconfidence. We investigate whether humans can learn to mentally recalibrate AI confidence signals through repeated experience. In a behavioral experiment (N = 200), participants predicted the AI's correctness across four AI calibration conditions: standard, overconfidence, underconfidence, and a counterintuitive "reverse confidence" mapping. Results demonstrate robust learning across all conditions, with participants significantly improving their accuracy, discrimination, and calibration alignment over 50 trials. We present a computational model utilizing a linear-in-log-odds (LLO) transformation and a Rescorla-Wagner learning rule to explain these dynamics. The model reveals that humans adapt by updating their baseline trust and confidence sensitivity, using asymmetric learning rates to prioritize the most informative errors. While humans can compensate for monotonic miscalibration, we identify a significant boundary in the reverse confidence scenario, where a substantial proportion of participants struggled to override initial inductive biases. These findings provide a mechanistic account of how humans adapt their trust in AI confidence signals through experience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。