提出多目标优化框架,同时提升神经网络模型的可解释性与公平性。
Interpretable and Fair Generalized Additive Neural Networks via Multi-objective Learning

- 基于多目标进化学习,联合优化准确率、可解释性与公平性。
- 揭示三者间复杂权衡关系,实证其内在机制。
- 支持实际部署的重训练策略,适合可信AI研究者使用。
可解释性与公平性是可信人工智能中备受关注的两大维度。现有研究多聚焦于提升基于神经网络的广义加性模型(NN-based GAMs)的准确性,而对其可解释性的系统性评估与优化仍不足。本文首次引入可量化的可解释性指标,实证其有效性,并提出基于多目标进化学习的多目标神经基模型(MONBM)框架,同时优化准确率、可解释性与公平性。进一步设计部分重训练策略,使进化多目标优化适用于深层模型架构。实验揭示了三者间的复杂权衡及其成因,表明多目标优化能有效结合自解释模型,揭示可信目标间的关联。所提方法生成一系列不同权衡的模型,性能优于当前主流方法。
原文摘要 · Abstract (English)
Interpretability and fairness are two of the most emphasized dimensions in trustworthy artificial intelligence (AI). Various explainable AI methods have been introduced to improve interpretability. This paper focuses on neural network (NN)-based generalized additive models (GAMs), a class of self-interpretable models. While most existing research has prioritized improving the accuracy of NN-based GAMs, their interpretability remains largely underexplored. To address this gap, this paper introduces explicit quantitative metrics for evaluating the interpretability of NN-based GAMs, empirically examines their effectiveness, and explores strategies for improving interpretability within these models. In addition, the simultaneous and explicit optimization of both interpretability and fairness, along with their trade-offs and the underlying reasons, remains underexplored. To address this, we propose a multi-objective neural basis model (MONBM) framework based on multi-objective evolutionary learning to consider accuracy, interpretability, and fairness simultaneously. A partial retraining strategy is further developed to facilitate the practical application of evolutionary multi-objective optimization to deep model architectures. Based on MONBM, this paper reveals the complex relationships between these dimensions and the reasons behind these intricate relationships. This analysis demonstrates how multi-objective optimization can be combined with self-interpretable models to reveal relationships among trustworthiness objectives. In addition, MONBM obtains a set of models with different trade-offs between dimensions, and the competitiveness of the approach is validated by comparing it with state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。