为可解释AI评估指标设计标准化报告卡片,提升透明度与可比性。
Evaluation Cards for XAI Metrics
- 提出XAI评估卡片模板,规范指标定义与报告
- 涵盖目标属性、假设条件、验证证据等核心内容
- 适合研究者和审稿人用于评估与对比方法
可解释AI(XAI)方法的评估因缺乏标准化而受限:指标定义不一致、报告不完整,且很少与通用基线进行验证。本文指出评估报告透明度是关键但被忽视的问题。我们提出XAI评估卡片,一种类比模型卡片的文档模板,旨在伴随任何引入XAI评估指标的研究。该卡片要求明确声明目标属性、可解释性层级、指标假设、验证证据、可规避风险及已知失效案例。我们认为,将此模板作为社区规范,可减少评估碎片化,支持元分析,并提升XAI研究的问责性。
原文摘要 · Abstract (English)
The evaluation of explainable AI (XAI) methods is affected by a lack of standardization. Metrics are inconsistently defined, incompletely reported, and rarely validated against common baselines. In this paper, we identify transparency of evaluation reporting as a central, under-addressed problem. We propose the XAI Evaluation Card, a documentation template analogous to model cards, designed to accompany any study that introduces an XAI evaluation metric. The card covers explicit declaration of target properties, grounding levels, metric assumptions, validation evidence, gaming risks, and known failure cases. We argue that adopting this template as a community norm would reduce evaluation fragmentation, support meta-analysis, and improve accountability in XAI research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。