用拉格朗日力学构建可解释模型的通用理论,实现可解释方法的逻辑推导设计。
The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

- 基于拉格朗日力学建立可解释性理论框架,从用户需求推导出对称性与约束。
- 通过优化目标函数最小值,获得最优可解释模型,解决现有方法局限。
- 为可解释性研究提供统一范式,适合研究人员与教学使用。
随着人工智能模型日益复杂,可解释性已成为理解、调试和控制其计算过程不可或缺的工具。然而,当前可解释性缺乏通用理论以逻辑推导设计可解释方法,导致文献碎片化与评估标准不一。为此,本文提出标准可解释模型(SIM),一个基于拉格朗日力学的通用理论,可系统推导可解释方法。SIM通过一组前提定义特定用户所需的可解释性,从中推导出可解释性对称性及对应约束,构成一个拉格朗日量,其极小值对应最优可解释模型。可通过调整黑箱模型参数使其更可解释,或直接将约束编译进可解释架构中。实验表明,SIM能识别并解决传统、概念型及机制型可解释方法的缺陷,揭示未充分探索的研究方向,并指导核心编程接口设计。除作为研究方法外,其演绎性质还为可解释性课程提供教学基础,有望重塑长期分散的该领域研究范式。
原文摘要 · Abstract (English)
As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations. However, interpretability lacks general theories to deductively design interpretable methods. This gap between theories and methods results in a fragmented literature and inconsistent evaluation protocols. To fill this gap, we introduce the Standard Interpretable Model (SIM), a general theory grounded in Lagrangian mechanics that enables the deductive design of interpretable methods. Specifically, the SIM summarises, in a set of premises, what interpretability is for a target user. From these premises, the SIM systematically derives interpretability symmetries and corresponding constraints, which shape the landscape of a Lagrangian whose minima correspond to optimal interpretable models. To reach the minima, one can either update the parameter values of an opaque model to make it more interpretable or compile constraints into an interpretable architecture. We empirically show that the SIM identifies and solves limitations of existing methods (including traditional, concept-based, and mechanistic interpretability), highlights underexplored research directions, and informs the design of core programming interfaces. Beyond being a research method, the deductive nature of the SIM offers pedagogical grounding for interpretability curricula and may shift the scientific community's perspective of a discipline that has long been fragmented.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。