厘清机器学习在物理中的可解释性与可解释性区别,强调二者是设计选择而非固有属性。
Interpreting "Interpretability" and Explaining "Explainability" in Machine Learning in Physics

- 区分模型结构透明性(可解释性)与科学内涵映射(可解释性)
- 指出可解释性与表达能力、可解释性与适应性间存在权衡
- 适合关注模型设计哲学与物理领域建模的科研人员
我们回顾了机器学习在物理领域中的可解释性与可解释性概念。将可解释性定义为模型结构透明度(理解或近似其内部机制的能力),将可解释性定义为模型的科学内容(将其映射到领域知识的能力)。讨论了两者带来的权衡(可解释性与表达能力;可解释性与适应性),以及各自适用的情境,并介绍了实现两者的内在与事后工具。始终强调,机器学习模型与经典模型面临相同的科学问题,仅规模不同;可解释性与可解释性应视为有意的设计选择,而非固有属性。同时强调任务定义与干预计划作为模型设计的核心要素。
原文摘要 · Abstract (English)
We review the concepts of interpretability and explainability as they apply to machine learning in physics. We define interpretability as concerning the structural transparency of a model (the ability to understand or approximate its inner workings) and explainability as concerning the scientific content of a model (the ability to map it onto domain knowledge). We discuss the trade-offs each entails (interpretability vs. expressivity; explainability vs. adaptability), the contexts in which each is needed, and the intrinsic and post-hoc tools available for achieving them. Throughout, we emphasize that machine-learned models are subject to the same scientific questions as classical models, differing only in scale, and that interpretability and explainability are best understood as deliberate modeling choices rather than inherent properties. We also emphasize the importance of task specification and intervention plans as a core aspect of model design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。