首个多层次因果干预框架,解析变分自编码器的生成机制。
A Multi-Level Causal Intervention Framework for Mechanistic Interpretability in Variational Autoencoders
- 设计四类干预操作,系统分析输入、隐空间与激活的因果影响。
- 发现因果效应强度与解耦度存在负相关,且β-VAE在复杂数据上性能下降。
- 揭示连续干预方法对离散隐变量表示不适用,适合研究生成模型机理者参考。
理解生成模型如何表征和转换数据是深度学习可解释性的基础问题。尽管判别模型的机制可解释性已取得显著进展,但针对变分自编码器(VAEs)的研究仍较少。本文提出首个通用的多层级因果干预框架,用于VAE的机制可解释性分析。该框架包含四种操作:输入操纵、隐空间扰动、激活补丁和因果中介分析。我们还定义了三项新量化指标:因果效应强度(CES)、干预特异性与电路模块化,这些指标捕捉了现有解耦度量无法涵盖的特性。我们在六种架构(标准VAE、beta-VAE、FactorVAE、beta-TC-VAE、DIP-VAE-II、VQ-VAE)和五个基准数据集(dSprites、3DShapes、MPI3D、CelebA、SmallNORB)上进行了迄今最大规模的实证研究,每配置运行三次,共90次独立训练。结果揭示:(i) 同一数据集内,CES与DCI解耦度呈稳定负相关(即CES-DCI权衡);(ii) beta-VAE的KL重加权机制在生成因子接近隐层维度时引发容量瓶颈,导致复杂数据上解耦度下降;(iii) 无单一架构在所有数据集上最优,最佳选择依赖于数据结构;(iv) 将CES指标应用于离散隐空间(如VQ-VAE)时,值趋近零,暴露出连续干预方法对离散表示的固有局限性。这些发现为生成模型的机制可解释性提供了理论基础与全面实证支持。
原文摘要 · Abstract (English)
Understanding how generative models represent and transform data is a foundational problem in deep learning interpretability. While mechanistic interpretability of discriminative architectures has yielded substantial insights, relatively little work has addressed variational autoencoders (VAEs). This paper presents the first general-purpose multilevel causal intervention framework for mechanistic interpretability of VAEs. The framework comprises four manipulation types: input manipulation, latent-space perturbation, activation patching, and causal mediation analysis. We also define three new quantitative metrics capturing properties not measured by existing disentanglement metrics alone: Causal Effect Strength (CES), intervention specificity, and circuit modularity. We conduct the largest empirical study to date of VAE causal mechanisms across six architectures (standard VAE, beta-VAE, FactorVAE, beta-TC-VAE, DIP-VAE-II, and VQ-VAE) and five benchmarks (dSprites, 3DShapes, MPI3D, CelebA, and SmallNORB), with three seeds per configuration, totaling 90 independent training runs. Our results reveal several findings: (i) a consistent within-dataset negative correlation between CES and DCI disentanglement (the CES-DCI trade-off); (ii) that the KL reweighting mechanism of beta-VAE induces a capacity bottleneck when generative factors approach latent dimensionality, degrading disentanglement on complex datasets; (iii) that no single VAE architecture dominates across all five datasets, with optimal choice depending on dataset structure; and (iv) that CES-based metrics applied to discrete latent spaces (VQ-VAE) yield near-zero values, revealing a critical limitation of continuous-intervention methods for discrete representations. These results provide both a theoretical foundation and comprehensive empirical evaluation for mechanistic interpretability of generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。