arXiv:2506.01059cs.LGcs.AI2025-06中稿 · FAccT 2025被引 4

用单元测试方式评测解释模型,发现不同方法的优劣。

XAI-Units: Benchmarking Explainability Methods with Unit Tests

  • 构建可复现的合成数据与模型,模拟各种推理行为
  • 揭示不同解释方法在特征交互等场景下的表现差异
  • 适合研究解释性方法的学者与开发可信AI系统的工程师

特征归因(FA)方法广泛用于可解释人工智能(XAI),帮助用户理解机器学习模型输入对输出的影响。然而,不同FA方法常对同一模型给出不一致的重要性评分。在缺乏真实标签或对模型内部机制了解不足的情况下,难以判断哪种方法在特定情境下更合适。为此,我们提出开源的XAI-Units基准,专门评估FA方法在多种模型行为(如特征交互、抵消效应、非连续输出)下的表现。该基准提供具有已知内在机制的配对数据集与模型,明确期望的归因分数。配套的内置评估指标使系统化实验更便捷,能清晰揭示FA方法在不同原子级模型推理模式中的表现。通过使用基于程序生成的模型与合成数据集,我们为客观、可靠的FA方法比较奠定了基础。

原文摘要 · Abstract (English)

Feature attribution (FA) methods are widely used in explainable AI (XAI) to help users understand how the inputs of a machine learning model contribute to its outputs. However, different FA models often provide disagreeing importance scores for the same model. In the absence of ground truth or in-depth knowledge about the inner workings of the model, it is often difficult to meaningfully determine which of the different FA methods produce more suitable explanations in different contexts. As a step towards addressing this issue, we introduce the open-source XAI-Units benchmark, specifically designed to evaluate FA methods against diverse types of model behaviours, such as feature interactions, cancellations, and discontinuous outputs. Our benchmark provides a set of paired datasets and models with known internal mechanisms, establishing clear expectations for desirable attribution scores. Accompanied by a suite of built-in evaluation metrics, XAI-Units streamlines systematic experimentation and reveals how FA methods perform against distinct, atomic kinds of model reasoning, similar to unit tests in software engineering. Crucially, by using procedurally generated models tied to synthetic datasets, we pave the way towards an objective and reliable comparison of FA methods.

可解释AI特征归因基准测试合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。