构建统一可解释性评分体系,评估AI模型的可信度。
Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models

- 从保真度、简洁性、稳定性三维度量化可解释性。
- 基于多数据集与模型实验建立离线知识库。
- 支持新场景下可解释性分数预测,适合模型开发者使用。
本文提出一个综合框架,用于评估LIME、SHAP等XAI方法在多个数据集和机器学习模型上的可解释性,目标是构建统一的多维可解释性评分。方法聚焦于保真度、简洁性和稳定性三个核心维度,通过基准测试系统评估并构建离线知识库,记录各注册模型的可解释性得分。该知识库可结合模型、数据集和XAI方法的元信息,对未见过的数据集与模型进行可解释性评分估计。保真度、简洁性与稳定性受数据集、模型及用户领域知识影响显著。通过三个开源数据集验证框架,分析结果揭示了不同数据特性对可解释性的影响。本工作为可解释人工智能(XAI)领域提供了一种稳健且通用的评估工具,有助于提升AI系统的透明性与可信度。
原文摘要 · Abstract (English)
In this paper, we present a comprehensive framework for assessing the explainability of various XAI methods, such as LIME and SHAP, across multiple datasets and machine learning models, with the ultimate goal of creating a unified multidimensional explainability score. Our methodology focuses on three key aspects of explainability: fidelity, simplicity, and stability. We leverage benchmarking experiments to systematically evaluate these aspects and use the insights gained to construct an offline knowledge base. This knowledge base captures the explainability scores for each registered model and serves as a valuable resource for context-dependent evaluation of explainability. By analyzing the complementary characteristics and metadata of AI models, datasets, and XAI methods, the knowledge base will enable the estimation of explainability scores for previously unseen datasets and models. Properties like fidelity, simplicity, and stability may vary significantly based on the dataset, underlying model, and domain expertise of the end user. We demonstrate our framework by applying it to three open-source datasets, discussing the implications of the obtained results in relation to the characteristics of the datasets. Our work contributes to the growing field of XAI by providing a robust and versatile tool for evaluating and comparing the explainability of various XAI methods, ultimately supporting the development of more transparent and trustworthy AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。