研究人如何决定何时信任AI,发现人类常错失正确建议或盲目相信错误答案。
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?

- 设计问答游戏实验,观察人类与AI协作中的委托与采纳决策
- 人类错失3.9%正确建议,1.7%时被错误AI误导,双方均出错
- 建议用可信度校准、证据解释和信任调节机制提升协作
AI系统存在缺陷,人类在判断是否信任AI时也可能犯错。因此,提升人机协作需理解人类在何时、为何以及如何选择依赖AI。本文研究两种独立的依赖行为:委托决策(决定何时让AI自主运行而不了解其输出)和采纳决策(评估AI建议后决定如何使用)。两者共同塑造协作效果,但以往研究很少在同一真实场景中同时考察。为此,我们设计了一个问答竞赛实验,23位专家人类与16个AI代理参与24场比赛,共记录387次委托和1440次采纳决策。结果显示,人机协作整体优于单独使用人类或AI,但人类仍做出非最优决策:错失3.9%的正确建议机会,且在AI误导时错误依赖率达1.7%。双方都会产生错误答案——当人类与AI意见不一致时,模型置信度接近随机水平;而当AI建议与人类初始错误答案一致时,确认偏误导致64.5%的过度依赖。为缩小差距,建议引入校准后的置信度、基于证据的解释,以及帮助用户调整信任的机制。
原文摘要 · Abstract (English)
AI systems are fallible, and humans can make mistakes in deciding whether to trust AI over their own judgment. Thus, improving human-AI collaboration requires understanding when, why, and how humans decide to rely on AI. We study two distinct reliance decisions: the delegation choice -- deciding when to let AI act autonomously without knowing its output, and the adoption choice -- evaluating AI suggestions and deciding how to use them. Both of these decoupled reliance patterns shape collaboration, but prior work rarely studies them together in realistic settings with the same users. We address this gap by studying collaborative human--AI teams competing in a question-answering game in which humans can choose when and how to work with AI agents to win. Our 24 matches pair 23 expert humans with 16 AI agents, capturing 387 delegation and 1440 adoption decisions. While human--AI collaboration performs better than either AI or humans alone, humans make suboptimal collaboration decisions, both under-relying on correct AI suggestions (3.9% of opportunities missed) and over-relying when AI misleads them (1.7%). Both parties contribute wrong answers: reported model confidence is near chance when humans and AI disagree, while confirmation bias drives higher under-reliance (64.5%) when an AI suggestion agrees with humans' initial incorrect answer. To close this gap, we recommend calibrated confidence, evidence-grounded explanations, and mechanisms that help users refine trust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。