arXiv:2603.25251cs.HCcs.AI2026-03

验证解释正确性对人类理解的影响,发现并非越准越好。

Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding

  • 用四种正确率(100%/85%/70%/55%)控制解释质量
  • 正确率低于70%时人类理解显著下降,但70%以下无进一步影响
  • 即使解释完全正确,也仅部分人能真正理解,需结合人类评估

解释型AI(XAI)方法通常通过功能正确性指标(如忠实度、保真度)评估,假设更高正确性带来更好人类理解。然而这一假设未在可控条件下验证。本研究开展用户实验(N=200),在合成时间序列分类任务中操控解释正确性至四个水平(100%、85%、70%、55%),参与者无法依赖领域知识或视觉直觉。解释正确性以已知真实情况为基准,而非模型估计。参与者根据特征归因类解释进行决策预测(前向模拟),其前向模拟准确率作为理解代理指标。结果显示:正确性影响理解,但非线性——70%与55%时准确率显著下降,70%以下无进一步降低;85%与100%间无显著差异。低正确性主要降低理解人数而非统一降低准确率。即使100%正确,理解仍呈双峰分布,仅部分人显著高于随机水平。探索性分析发现,部分人预测正确却误述模型依据;自评评分仅在完全正确时与前向模拟准确率相关。结果表明,功能正确性差异并不总对应人类理解差异,强调需将功能性度量与人类表现挂钩。

原文摘要 · Abstract (English)

Explainable AI (XAI) methods are commonly evaluated using functional correctness metrics, sometimes termed faithfulness or fidelity, which estimate how closely an explanation reflects the model's reasoning. Higher correctness is assumed to produce better human understanding, but this link has not been tested with controlled levels. We conducted a user study (N=200) that manipulated explanation correctness at four levels (100%, 85%, 70%, 55%) in a synthetic time series classification task where participants could not rely on domain knowledge or visual intuition. Correctness was defined against a known ground truth, not estimated from a trained model. Participants predicted a simulated AI's decisions from feature-attribution-style explanations (forward simulation), and we used their forward simulation accuracy as a proxy for understanding. Correctness affected understanding, but not at every level: forward simulation accuracy dropped at 70% and 55% relative to fully correct explanations, with no further reduction below 70% and no conclusive difference between 85% and 100%. Lower correctness reduced how many participants predicted accurately rather than lowering accuracy uniformly, and even fully correct explanations did not guarantee understanding: forward simulation accuracy was bimodal there, with only a subset of participants well above chance. In exploratory analyses, some participants predicted decisions accurately but described the wrong pattern when asked what the AI used, and self-reported ratings correlated with forward simulation accuracy only when explanations were fully correct. These findings show that not all differences in functional correctness translate to differences in human understanding, highlighting the need to validate functional metrics against human outcomes.

可解释AI人类评估忠实度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。