通过调试AI系统来训练用户合理信任,结果发现反而降低了信任度。
To Err Is AI! Debugging as an Intervention to Facilitate Appropriate Reliance on AI Systems
- 用调试AI作为干预手段,检验能否提升用户对AI的信任判断
- 234名参与者实验显示,调试后用户对AI依赖度下降
- 暴露AI弱点可能引发信心动摇,适合研究人机信任机制者参考
强大的预测型AI系统在辅助人类决策方面展现出巨大潜力。近期实证研究指出,理想的人机协作需建立在‘恰当信任’之上。然而,在缺乏针对具体实例的AI表现反馈时,准确评估其可信度极具挑战性,尤其当模型在分布外数据上表现差异大时,基于特定数据集的反馈不可靠。受批判性思维文献启发,本文提出将调试AI系统作为一种干预手段,以促进用户形成恰当依赖。通过一项定量实证研究(N = 234),我们发现该调试干预并未如预期般提升恰当依赖,反而导致用户对AI系统的依赖程度下降——可能是由于早期暴露于AI的缺陷所致。我们进一步分析了不同性能组别中用户信心与对AI可信度估计的变化动态,揭示不恰当依赖模式的形成机制。研究结果对设计有效干预策略、优化人机协作具有重要启示。
原文摘要 · Abstract (English)
Powerful predictive AI systems have demonstrated great potential in augmenting human decision making. Recent empirical work has argued that the vision for optimal human-AI collaboration requires 'appropriate reliance' of humans on AI systems. However, accurately estimating the trustworthiness of AI advice at the instance level is quite challenging, especially in the absence of performance feedback pertaining to the AI system. In practice, the performance disparity of machine learning models on out-of-distribution data makes the dataset-specific performance feedback unreliable in human-AI collaboration. Inspired by existing literature on critical thinking and a critical mindset, we propose the use of debugging an AI system as an intervention to foster appropriate reliance. In this paper, we explore whether a critical evaluation of AI performance within a debugging setting can better calibrate users' assessment of an AI system and lead to more appropriate reliance. Through a quantitative empirical study (N = 234), we found that our proposed debugging intervention does not work as expected in facilitating appropriate reliance. Instead, we observe a decrease in reliance on the AI system after the intervention -- potentially resulting from an early exposure to the AI system's weakness. We explore the dynamics of user confidence and user estimation of AI trustworthiness across groups with different performance levels to help explain how inappropriate reliance patterns occur. Our findings have important implications for designing effective interventions to facilitate appropriate reliance and better human-AI collaboration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。