从分布视角重构视觉机制可解释性,平衡人类感知与模型忠实性。
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
- 基于KL最小化构建分布视角的可解释性框架
- 揭示传统方法在感知与模型忠实性上的双重偏差
- 采用能量引导扩散采样实现更可靠的解释
当前视觉机制可解释性(MI)方法多依赖启发式手段(如前K激活检索或正则化优化),难以保证解释的有效性。本文提出一种分布视角的理论框架,将特征激活对自然图像分布的影响建模为KL最小化优化问题。该框架揭示了现有方法存在两大缺陷:一是解释结果偏离自然图像分布,导致人类难以理解;二是无法有效激活模型内部特征,违背模型机制。为此,本文提出基于KL最小软约束原则的新模型,通过能量引导的扩散后验采样实现理论上的可解释性与忠实性平衡。大量实验验证了该框架的理论合理性,并在DINOv3视觉模型上展现了实际有效性。
原文摘要 · Abstract (English)
Most current paradigms in visual mechanistic interpretability (MI) remain confined to interpreting internal units of the vision model via heuristic methods (e.g., top-$K$ activation retrieval or optimization with regularization). In this work, we establish a theoretical distributional view for visual MI, which models the influence of a feature activation on the natural image distribution, thereby formulating a Kullback-Leibler (KL)-minimal optimization problem to model the MI task. Under this framework, statistical biases are identified within previous MI paradigms, which reveal that they may either be perceptually uninterpretable to humans (i.e., deviate from the natural image distribution), or mechanistically unfaithful to the vision models (i.e., unable to activate model features). To resolve the biases under the distributional view, we propose a model with a KL-minimal soft-constraint principle for visual MI that theoretically balances interpretability and faithfulness. We realize this principle via energy-guided diffusion posterior sampling. Extensive experiments validate the theoretical soundness of the proposed distributional view and demonstrate the practical effectiveness of our paradigm on the DINOv3 vision model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。