系统梳理自解释神经网络的五大方法,助力模型决策透明化。
A Comprehensive Survey on Self-Interpretable Neural Networks
- 从五类视角归纳自解释模型设计思路:基于属性、函数、概念、原型和规则。
- 提供图像、文本、图数据等多场景可视化解释案例,验证方法适用性。
- 汇总评估指标与开放挑战,适合关注AI可解释性的研究者参考。
神经网络在多个领域取得显著成果,但缺乏可解释性限制了其在关键决策场景中的应用。事后解释方法常面临鲁棒性和保真度问题,推动了自解释神经网络的发展——这类模型通过结构设计内生揭示预测依据。尽管已有事后解释的综述,但针对自解释神经网络的系统性总结仍不足。本文首次全面梳理该领域工作,从五个核心维度(基于属性、函数、概念、原型、规则)系统归纳方法,并展示图像、文本、图数据及深度强化学习等场景下的可视化解释实例。同时,总结现有可解释性评估指标,指出当前挑战并提出未来方向。为支持持续发展,作者提供公开资源库:https://github.com/yangji721/Awesome-Self-Interpretable-Neural-Network。
原文摘要 · Abstract (English)
Neural networks have achieved remarkable success across various fields. However, the lack of interpretability limits their practical use, particularly in critical decision-making scenarios. Post-hoc interpretability, which provides explanations for pre-trained models, is often at risk of robustness and fidelity. This has inspired a rising interest in self-interpretable neural networks, which inherently reveal the prediction rationale through the model structures. Although there exist surveys on post-hoc interpretability, a comprehensive and systematic survey of self-interpretable neural networks is still missing. To address this gap, we first collect and review existing works on self-interpretable neural networks and provide a structured summary of their methodologies from five key perspectives: attribution-based, function-based, concept-based, prototype-based, and rule-based self-interpretation. We also present concrete, visualized examples of model explanations and discuss their applicability across diverse scenarios, including image, text, graph data, and deep reinforcement learning. Additionally, we summarize existing evaluation metrics for self-interpretability and identify open challenges in this field, offering insights for future research. To support ongoing developments, we present a publicly accessible resource to track advancements in this domain: https://github.com/yangji721/Awesome-Self-Interpretable-Neural-Network.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。