arXiv:2506.22740cs.AIstat.ML2025-06

用决策理论评估解释效果,看它到底帮人做对了什么。

Explanations are a Means to an End: Decision Theoretic Explanation Evaluation

  • 把解释看作能提升决策的信号,从理论上评估其价值上限。
  • 发现人类已有判断已吸收部分解释价值,剩余潜力可量化。
  • 适用于人机协作决策与模型可解释性研究,指导解释设计。

模型解释通常通过与实际用途关联较弱的代理指标进行评估。本文提出一种决策理论框架,将解释视为能提升特定决策任务表现的信息信号。该框架定义了三个可量化的评估目标:1)任何获得解释的主体所能达到的最佳性能上限;2)在现有人类决策策略基础上,解释所能带来的额外理论价值;3)向人类决策者提供解释所产生的因果效应。我们设计了可操作的验证流程,并应用于人机决策支持与机制可解释性研究中,评估解释潜力并分析人类行为变化。

原文摘要 · Abstract (English)

Explanations of model behavior are commonly evaluated via proxy properties weakly tied to the purposes explanations serve in practice. We contribute a decision theoretic framework that treats explanations as information signals valued by the expected improvement they enable on a specified decision task. This approach yields three distinct estimands: 1) a theoretical benchmark that upperbounds achievable performance by any agent with the explanation, 2) a human-complementary value that quantifies the theoretically attainable value that is not already captured by a baseline human decision policy, and 3) a behavioral value representing the causal effect of providing the explanation to human decision-makers. We instantiate these definitions in a practical validation workflow, and apply them to assess explanation potential and interpret behavioral effects in human-AI decision support and mechanistic interpretability.

解释评估决策理论人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。