arXiv:2505.06258cs.LGcs.AI2025-05

统一解释框架ABE提升模型可解释性,支持自定义开发与验证。

ABE: A Unified Framework for Robust and Faithful Attribution-Based Explainability

  • 提出统一框架ABE,整合多种归因方法并满足理论公理
  • 四模块设计实现可扩展的鲁棒性、可解释性与验证能力
  • 适合研究可解释性算法的开发者与需要透明AI的实践者

归因算法对于提升深度学习模型的可解释性与可信度至关重要,能识别影响模型决策的关键特征。现有框架如InterpretDL和OmniXAI虽集成多种归因方法,但存在可扩展性差、耦合度高、理论约束强及用户友好性不足等问题,阻碍了神经网络的透明性与互操作性。为此,我们提出归因式可解释性(Attribution-Based Explainability, ABE),一个形式化基础归因方法并集成前沿归因算法的统一框架,确保符合归因公理。ABE支持研究人员开发新归因技术,并通过四个可定制模块——鲁棒性、可解释性、验证、数据与模型——增强可解释性。该框架为推进基于归因的可解释性提供可扩展、可扩展的基础,助力透明AI系统建设。代码已开源:https://github.com/LMBTough/ABE-XAI。

原文摘要 · Abstract (English)

Attribution algorithms are essential for enhancing the interpretability and trustworthiness of deep learning models by identifying key features driving model decisions. Existing frameworks, such as InterpretDL and OmniXAI, integrate multiple attribution methods but suffer from scalability limitations, high coupling, theoretical constraints, and lack of user-friendly implementations, hindering neural network transparency and interoperability. To address these challenges, we propose Attribution-Based Explainability (ABE), a unified framework that formalizes Fundamental Attribution Methods and integrates state-of-the-art attribution algorithms while ensuring compliance with attribution axioms. ABE enables researchers to develop novel attribution techniques and enhances interpretability through four customizable modules: Robustness, Interpretability, Validation, and Data & Model. This framework provides a scalable, extensible foundation for advancing attribution-based explainability and fostering transparent AI systems. Our code is available at: https://github.com/LMBTough/ABE-XAI.

可解释性归因方法AI透明化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。