arXiv:2607.08641cs.LG2026-07

用部分依赖约束引导神经网络训练,让模型解释更符合领域知识。

Steering Neural Network Training through Interpretable Constraints Based on Partial Dependence

  • 通过部分依赖设计可解释的约束条件,引导模型学习特定特征响应。
  • 在动态系统预测等回归任务中,模型性能提升且数据效率更高。
  • 适合需模型解释与领域知识对齐的研究者,尤其关注可解释性的人工智能应用。

近年来,机器学习可解释性受到广泛关注。尽管已有大量工作致力于解析模型学习到的特征交互,但评估解释质量的研究仍较少,更少有研究关注如何调整模型使其解释符合先验知识,即解释引导学习。现有方法多聚焦分类问题,且常假设已知关键特征或区域。本文提出一种基于部分依赖的神经网络引导训练新方法,使模型对特定特征的平均响应与具体问题的领域知识一致。我们在多个回归任务(包括动态系统预测)上实证验证:经该方法控制训练的模型性能优于无约束模型,且更具数据效率。此外,前者的解释结果与用户提供的知识相符,后者则不一致。

原文摘要 · Abstract (English)

Over the last few years, there has been an increased interest in making machine learning models more interpretable. Although a great deal of effort goes into developing techniques for interpreting the interactions learned by a given model, fewer studies focus on assessing the quality of such explanations. Even fewer focus on how to adjust the model to produce explanations faithful to prior knowledge, a process known as explanation-guided learning. Furthermore, most approaches in this area focus on classification problems and usually assume prior knowledge about which input features or regions are most important. In this work, we introduce a new approach to steering neural networks based on partial dependence, such that their average response to certain features aligns with specific functional domain knowledge about the problem. We empirically demonstrate on a range of regression problems, including dynamical systems forecasting, that models whose training has been controlled using our method perform better than unconstrained models and are more data-efficient. Moreover, we highlight that interpretations obtained from the former actually align with the user-provided knowledge, whereas those obtained from the latter do not.

可解释性神经网络部分依赖领域知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。