研究生成式AI在不同难度任务中的错误如何影响用户依赖度。
Effects of Generative AI Errors on User Reliance Across Task Difficulty
- 通过图表生成任务实验,人为控制AI错误率
- 简单任务出错时用户仍保持较高依赖,不比复杂任务更退缩
- 揭示用户对AI错误模式的容忍度,适合人机交互研究者参考
人工智能的能力呈现不规则边界:某些人类轻松的任务,AI反而失败;而人类困难的任务,AI却可能成功。为探究用户对这一现象的反应,我们设计了一种激励相容的实验方法,基于图表生成任务,人为引入生成式AI的错误,并测试其对用户依赖度的影响。我们在一项预注册的3×2实验中(N = 577)设置了10%、30%或50%的错误率,分别作用于简单或困难的图表生成任务。结果表明,错误越多,用户使用越少;但令人意外的是,简单任务中的错误并未显著降低用户依赖度,高于复杂任务中的错误影响,暗示在该实验设置下,用户并不排斥这种非线性的错误模式。我们呼吁未来研究同时改变任务难度与错误特征(如错误模式是否易学习)以深入探索。
原文摘要 · Abstract (English)
The capabilities of artificial intelligence (AI) lie along a jagged frontier, where AI systems surprisingly fail on tasks that humans find easy and succeed on tasks that humans find hard. To investigate user reactions to this phenomenon, we developed an incentive-compatible experimental methodology based on diagram generation tasks, in which we induce errors in generative AI output and test effects on user reliance. We demonstrate the interface in a preregistered 3x2 experiment (N = 577) with error rates of 10%, 30%, or 50% on easier or harder diagram generation tasks. We confirmed that observing more errors reduces use, but we unexpectedly found that easy-task errors did not significantly reduce use more than hard-task errors, suggesting that people are not averse to jaggedness in this experimental setting. We encourage future work that varies task difficulty at the same time as other features of AI errors, such as whether the jagged error patterns are easily learned.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。