提出固定点解释新方法,让模型解释更稳定可靠
Fixed Point Explainability
- 基于递归验证机制,评估解释的稳定性与忠实性
- 在Llama-3.3-70B等模型上验证了解释收敛性
- 适合研究模型可解释性与潜在缺陷的学者
本文引入一种受“为何回归”原则启发的固定点解释形式化概念,通过递归应用评估模型与其解释器之间相互作用的稳定性。固定点解释具备最小性、稳定性和忠实性等性质,能揭示隐藏的模型行为和解释弱点。我们为多种解释器类别(从特征基到机制型如稀疏自编码器)定义了收敛条件,并在多个数据集和模型上报告了定量与定性结果,包括Llama-3.3-70B等大语言模型。
原文摘要 · Abstract (English)
This paper introduces a formal notion of fixed point explanations, inspired by the "why regress" principle, to assess, through recursive applications, the stability of the interplay between a model and its explainer. Fixed point explanations satisfy properties like minimality, stability, and faithfulness, revealing hidden model behaviours and explanatory weaknesses. We define convergence conditions for several classes of explainers, from feature-based to mechanistic tools like Sparse AutoEncoders, and we report quantitative and qualitative results for several datasets and models, including LLMs such as Llama-3.3-70B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。