用对称性定义可行动的可解释性,让模型设计与验证有据可依。
Actionable Interpretability Must Be Defined in Terms of Symmetries
- 以四种对称性为基石,构建可解释模型的数学框架。
- 将推理、干预、反事实等统一为贝叶斯逆问题。
- 为安全合规验证提供可测试的形式化工具,适合安全敏感场景。
本文认为当前人工智能可解释性研究本质上定义不清,因现有定义无法形式化检验或设计可解释性。我们提出,可行动的可解释性必须基于对称性来定义,这些对称性可指导模型设计并生成可测试条件。在概率视角下,我们假设四种对称性——推理等变性、信息不变性、概念闭包不变性与结构不变性——足以实现三方面目标:(i) 将可解释模型形式化为概率模型的一个子类;(ii) 统一表述可解释推理(如对齐、干预与反事实)为贝叶斯逆过程;(iii) 提供形式化框架,用于验证模型是否符合安全标准与法规要求。
原文摘要 · Abstract (English)
This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpretability can be formally tested or designed for. We posit that actionable definitions of interpretability must be formulated in terms of *symmetries* that inform model design and lead to testable conditions. Under a probabilistic view, we hypothesise that four symmetries (inference equivariance, information invariance, concept-closure invariance, and structural invariance) suffice to (i) formalise interpretable models as a subclass of probabilistic models, (ii) yield a unified formulation of interpretable inference (e.g., alignment, interventions, and counterfactuals) as a form of Bayesian inversion, and (iii) provide a formal framework to verify compliance with safety standards and regulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。