arXiv:2505.07910cs.LGcs.AI2025-05

让神经网络在性能与解释一致性间取得平衡,提升可信度。

Tuning for Trustworthiness -- Balancing Performance and Explanation Consistency in Neural Network Optimization

  • 引入解释一致性概念,量化不同归因方法的一致性
  • 首次将解释一致性纳入超参数优化目标,实现多目标权衡
  • 发现性能与解释力的折中区域,适合追求可靠模型的场景

尽管可解释人工智能(XAI)研究日益兴起,但超参数调优和神经架构优化仍主要关注最小化预测损失,极少考虑解释性。本文提出全新的XAI一致性概念,即不同特征归因方法之间的共识,并构建新指标进行量化。首次将XAI一致性直接融入超参数优化目标,建立兼顾预测性能与解释鲁棒性的多目标优化框架。基于顺序参数优化工具箱(SPOT),采用加权聚合与效用导向策略指导模型选择。通过该框架及配套工具,我们揭示了架构配置空间中的三类区域:性能差且可解释性低的区域,性能强但解释性弱(因低XAI一致性)的区域,以及在性能与可解释性间取得平衡的折中区域。本研究为未来探索折中区模型是否因避免过拟合训练性能而具备更强鲁棒性、更可靠外分布预测提供了基础。

原文摘要 · Abstract (English)

Despite the growing interest in Explainable Artificial Intelligence (XAI), explainability is rarely considered during hyperparameter tuning or neural architecture optimization, where the focus remains primarily on minimizing predictive loss. In this work, we introduce the novel concept of XAI consistency, defined as the agreement among different feature attribution methods, and propose new metrics to quantify it. For the first time, we integrate XAI consistency directly into the hyperparameter tuning objective, creating a multi-objective optimization framework that balances predictive performance with explanation robustness. Implemented within the Sequential Parameter Optimization Toolbox (SPOT), our approach uses both weighted aggregation and desirability-based strategies to guide model selection. Through our proposed framework and supporting tools, we explore the impact of incorporating XAI consistency into the optimization process. This enables us to characterize distinct regions in the architecture configuration space: one region with poor performance and comparatively low interpretability, another with strong predictive performance but weak interpretability due to low \gls{xai} consistency, and a trade-off region that balances both objectives by offering high interpretability alongside competitive performance. Beyond introducing this novel approach, our research provides a foundation for future investigations into whether models from the trade-off zone-balancing performance loss and XAI consistency-exhibit greater robustness by avoiding overfitting to training performance, thereby leading to more reliable predictions on out-of-distribution data.

可解释性超参数优化神经网络可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。