让模型既准又可解释,用图结构引导训练过程。
Interpretability-Guided Bi-objective Optimization: Aligning Accuracy and Explainability
- 用有向无环图建模特征重要性层级,结合中心极限定理保证结构合理性。
- 提出相对重要性得分,量化特征随时间累积的贡献度,提升可解释性。
- 兼顾准确率与可解释性的双目标优化,适合需要透明决策的高风险场景。
本文提出解释性引导的双目标优化(IGBO)框架,通过双目标形式将结构化领域知识融入可解释模型的训练过程。IGBO利用基于中心极限定理的方法构建特征重要性层次的有向无环图(DAG),并采用时序集成梯度(TIG)衡量特征重要性。提出新的相对重要性得分 Hk(X, θ),用于量化每个特征随时间累积的归因贡献。设计几何投影映射 P 以融合任务与可解释性梯度,并证明其收敛至帕累托驻点。针对 TIG 计算中的分布外问题,提出最优路径预言机架构,暂留作未来工作。基于中心极限定理构造的可解释性 DAG 提供了无条件的中位数阈值保证及高置信度下的条件保证,确保图的无环性与传递性。
原文摘要 · Abstract (English)
This paper introduces Interpretability-Guided Bi-objective Optimization (IGBO), a framework that trains interpretable models by incorporating structured domain knowledge via a bi-objective formulation. IGBO encodes feature importance hierarchies as a Directed Acyclic Graph (DAG) via Central Limit Theorem-based construction and uses Temporal Integrated Gradients (TIG) to measure feature importance. The framework employs a novel Relative Importance Score Hk(X, θ) that quantifies the normalized cumulative attribution of each feature over time. We propose a geometric projection mapping P for combining task and interpretability gradients, and prove convergence to Pareto-stationary points. To address the Out-of-Distribution problem in TIG computation, we outline an Optimal Path Oracle architecture, which we leave for future work. Central Limit Theorem-based construction of the interpretability DAG provides statistical guarantees on acyclicity and transitivity, with an unconditional guarantee for the median threshold and conditional guarantees for higher confidence levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。