对比多种原型网络模型,评估其可解释性与性能表现。
Comprehensive Evaluation of Prototype Neural Networks
- 构建多维度评估体系,涵盖标准与新提出的可解释性指标。
- 在细粒度分类、非独立同分布及多标签数据集上验证模型表现。
- 开源代码库支持快速应用与扩展新模型、新指标。
原型模型是可解释人工智能(XAI)和可解释机器学习的重要方法。本文对一系列代表性原型模型(包括ProtoPNet、ProtoPool和PIPNet)进行了深入分析。评估过程中,我们采用全面的度量指标,不仅包含文献中的标准指标,还提出了若干新指标以进一步完善模型可解释性的分析。实验在多样化的数据集上展开,涵盖细粒度分类、非独立同分布(Non-IID)设置及多标签分类任务,以更全面地对比各模型性能。此外,我们还开源了代码库(https://github.com/uos-sis/quanproto),支持指标的简易调用与扩展,便于用户轻松添加新指标或新模型。
原文摘要 · Abstract (English)
Prototype models are an important method for explainable artificial intelligence (XAI) and interpretable machine learning. In this paper, we perform an in-depth analysis of a set of prominent prototype models including ProtoPNet, ProtoPool and PIPNet. For their assessment, we apply a comprehensive set of metrics. In addition to applying standard metrics from literature, we propose several new metrics to further complement the analysis of model interpretability. In our experimentation, we apply the set of prototype models on a diverse set of datasets including fine-grained classification, Non-IID settings and multi-label classification to further contrast the performance. Furthermore, we also provide our code as an open-source library (https://github.com/uos-sis/quanproto), which facilitates simple application of the metrics itself, as well as extensibility -- providing the option for easily adding new metrics and models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。