arXiv:2605.08934cs.LG2026-05被引 1

用数学框架让神经网络解释更可验证、可组合、更易懂。

From Mechanistic to Compositional Interpretability

论文配图:From Mechanistic to Compositional Interpretability
图 1 · 摘自论文原文
  • 基于范畴论构建可组合的解释框架,确保结构与行为一致。
  • 将解释质量拆解为忠实度与复杂度,实现可优化的解释目标。
  • 提出压缩精炼方法,自动简化模型结构而不改变功能。

机械可解释性旨在通过逆向工程将神经模型的计算结构转化为人类可理解的组件。然而缺乏形式化框架时,机械解释无法客观验证、比较或组合。本文提出组合可解释性,一种基于组合性与最小描述长度原则的范畴论框架。组合解释由语法与语义映射构成,必须满足交换性以保证模型分解与行为的一致性。我们将解释质量分解为忠实度与复杂度,将可解释性建模为受约束的优化问题,并引入压缩精炼方法,系统重构模型为更简单的部分而保持功能不变。最后,我们推导出一个简约准则,在该准则下,语法压缩理论上能保证更简洁、更符合人类认知的解释。本框架将主流机械方法视为精炼的子类,并阐明其可压缩性启发式为何常与人类可解释性对齐。工作提供了一个可测量、可优化的自动化发现与评估机械解释蓝图。

原文摘要 · Abstract (English)

Mechanistic interpretability aims to explain neural model behaviour by reverse-engineering learned computational structure into human-understandable components. Without a formal framework, however, mechanistic explanations cannot be objectively verified, compared, or composed. We introduce compositional interpretability, a category-theoretic framework grounded in the principles of compositionality and minimum description length. Compositional interpretations are pairs of syntactic and semantic mappings that must commute to enforce consistency between a model's decomposition and its observed behaviour. We deconstruct explanation quality into measures of faithfulness and complexity to cast interpretability as a constrained optimisation problem, and introduce compressive refinement to systematically restructure models into simpler parts without altering their function. Finally, we derive a parsimony criterion under which syntactic compression theoretically guarantees more concise, human-aligned explanations. Our framework situates prominent mechanistic methods as subclasses of refinement, and clarifies why their compressibility heuristics tend to align with human interpretability. Our work provides a measurable, optimisable blueprint for automating the discovery and evaluation of mechanistic explanations.

可解释性范畴论模型压缩机制解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。