arXiv:2511.08361cs.LG2025-11被引 2

提出ProtoScore框架,统一评估原型类可解释AI方法在时序数据上的表现。

From Confusion to Clarity: ProtoScore -- A Framework for Evaluating Prototype-Based XAI

  • 基于Nauta等人提出的Co-12属性构建评估体系
  • 支持跨数据类型比较原型解释方法,尤其聚焦时序数据
  • 提供客观基准,避免依赖主观用户研究,适合实践者选型

神经网络的复杂性与不透明性在医疗、金融、法律等高风险领域带来挑战,理解其决策过程至关重要。可解释人工智能(XAI)致力于揭示模型决策逻辑,以建立合理信任并验证结果公平性。其中,基于原型的解释方法通过代表性样本阐明模型行为,展现出良好前景。然而,缺乏标准化评估基准,尤其在时间序列数据上,导致评价主观化,阻碍领域进展。本文提出ProtoScore框架,旨在对不同数据类型的原型类XAI方法进行可靠评估,重点覆盖时序数据。该框架整合Nauta等人提出的Co-12属性,实现原型方法间及与其他XAI方法间的有效对比,帮助实践者选择合适解释方法,降低用户研究成本。所有代码已公开于https://github.com/HelenaM23/ProtoScore。

原文摘要 · Abstract (English)

The complexity and opacity of neural networks (NNs) pose significant challenges, particularly in high-stakes fields such as healthcare, finance, and law, where understanding decision-making processes is crucial. To address these issues, the field of explainable artificial intelligence (XAI) has developed various methods aimed at clarifying AI decision-making, thereby facilitating appropriate trust and validating the fairness of outcomes. Among these methods, prototype-based explanations offer a promising approach that uses representative examples to elucidate model behavior. However, a critical gap exists regarding standardized benchmarks to objectively compare prototype-based XAI methods, especially in the context of time series data. This lack of reliable benchmarks results in subjective evaluations, hindering progress in the field. We aim to establish a robust framework, ProtoScore, for assessing prototype-based XAI methods across different data types with a focus on time series data, facilitating fair and comprehensive evaluations. By integrating the Co-12 properties of Nauta et al., this framework allows for effectively comparing prototype methods against each other and against other XAI methods, ultimately assisting practitioners in selecting appropriate explanation methods while minimizing the costs associated with user studies. All code is publicly available at https://github.com/HelenaM23/ProtoScore .

可解释AI原型方法时序数据评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。