提出量化人类对AI预测错误程度的新方法,提升可解释AI研究的严谨性。
How to Measure Human-AI Prediction Accuracy in Explainable AI Systems
- 设计三种数学框架衡量预测错误的细微差别
- 在36动作空间中验证方法,显著优于传统二元判断
- 适用于大输出空间的AI系统用户研究,提升实验可信度
评估可解释AI系统行为常依赖人类预测其下一步决策,但传统二元判断(对/错)在动作空间增大时失效,因正确率过低导致地板效应。本文提出三种数学基础来度量‘部分错误’的程度,并在两个序列决策场景中验证:一是实验室研究,86名参与者面对36个动作的选择;二是对先前4动作空间研究的再分析。该任务操作化与分析方法能显著提升大输出空间场景下用户研究的严谨性,为后续研究提供可复现的评估范式。
原文摘要 · Abstract (English)
Assessing an AI system's behavior-particularly in Explainable AI Systems-is sometimes done empirically, by measuring people's abilities to predict the agent's next move-but how to perform such measurements? In empirical studies with humans, an obvious approach is to frame the task as binary (i.e., prediction is either right or wrong), but this does not scale. As output spaces increase, so do floor effects, because the ratio of right answers to wrong answers quickly becomes very small. The crux of the problem is that the binary framing is failing to capture the nuances of the different degrees of "wrongness." To address this, we begin by proposing three mathematical bases upon which to measure "partial wrongness." We then uses these bases to perform two analyses on sequential decision-making domains: the first is an in-lab study with 86 participants on a size-36 action space; the second is a re-analysis of a prior study on a size-4 action space. Other researchers adopting our operationalization of the prediction task and analysis methodology will improve the rigor of user studies conducted with that task, which is particularly important when the domain features a large output space.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。