arXiv:2502.06775cs.LG2025-02ICML被引 1

通过约束概念表示优化,提升可解释模型的准确率与效率

Enhancing Performance of Explainable AI Models with Constrained Concept Refinement

  • 在保持可解释性的前提下,约束优化概念嵌入以减少性能损失
  • 算法实现零损失,同时逐步增强模型可解释性
  • 在多基准测试中显著提升准确率且计算成本更低

可解释性与准确率之间的权衡是机器学习长期存在的挑战,尤其对旨在实现可信可解释性的新型可解释设计方法而言更为突出,这类方法常因牺牲准确率而受限。本文研究了概念表示偏差对可解释模型预测性能的影响,提出一种新框架以缓解此问题。该框架基于在保留可解释性约束下的概念嵌入优化原则。以生成模型为测试平台,我们严格证明所提算法可在保持零损失的同时逐步提升模型可解释性。此外,我们在多个图像分类基准上评估了该框架在生成可解释预测方面的实际表现。相较于现有可解释方法,本方法不仅在多个大规模基准上同时提升预测准确率并保持可解释性,还实现了显著更低的计算开销。

原文摘要 · Abstract (English)

The trade-off between accuracy and interpretability has long been a challenge in machine learning (ML). This tension is particularly significant for emerging interpretable-by-design methods, which aim to redesign ML algorithms for trustworthy interpretability but often sacrifice accuracy in the process. In this paper, we address this gap by investigating the impact of deviations in concept representations-an essential component of interpretable models-on prediction performance and propose a novel framework to mitigate these effects. The framework builds on the principle of optimizing concept embeddings under constraints that preserve interpretability. Using a generative model as a test-bed, we rigorously prove that our algorithm achieves zero loss while progressively enhancing the interpretability of the resulting model. Additionally, we evaluate the practical performance of our proposed framework in generating explainable predictions for image classification tasks across various benchmarks. Compared to existing explainable methods, our approach not only improves prediction accuracy while preserving model interpretability across various large-scale benchmarks but also achieves this with significantly lower computational cost.

可解释AI概念表示模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。