arXiv:2603.21236cs.LG2026-03

发现图像模型的可解释性机制在表格数据上失效,提出新方法提升可解释性。

Posterior-Calibrated Causal Circuits in Variational Autoencoders: Why Image-Domain Interpretability Fails on Tabular Data

  • 引入后验校准的因果效应强度评估,改进表格数据的可解释性分析。
  • 表格VAE的模块化程度仅为图像模型的50%,且β-VAE因果效应几乎崩溃。
  • 新方法能有效识别架构差异,适合用于数据清洗与异常检测任务。

尽管基于机制的可解释性已为判别网络带来深刻洞察,但生成模型尤其是非图像领域的理解仍不充分。本文探究图像领域变分自编码器(VAE)中的因果结构能否推广至表格数据,因VAE正被广泛用于缺失值填补、异常检测和合成数据生成。我们扩展四层因果干预框架,覆盖五个不同VAE架构下的四组表格与一组图像基准,共完成75次训练运行及每轮三次随机种子实验。提出三种新方法:后验校准因果效应强度(CES)、路径特定激活修补和特征组解纠缠(FGD)。实验表明:(i)表格型VAE的模块化程度约为图像模型的50%;(ii)β-VAE在异构表格特征上因果效应近乎崩溃(表数据CES=0.043,图像数据=0.133),这与重构质量下降高度相关(相关系数r = -0.886);(iii)CES成功捕捉十一项中九项统计显著的架构差异(采用Holm–Šidák校正);(iv)高特异性干预预测出最高下游AUC值(r = 0.460,p < .001)。研究挑战了将图像研究经验直接迁移至表格数据的普遍假设。

原文摘要 · Abstract (English)

Although mechanism-based interpretability has generated an abundance of insight for discriminative network analysis, generative models are less understood -- particularly outside of image-related applications. We investigate how much of the causal circuitry found within image-related variational autoencoders (VAEs) will generalize to tabular data, as VAEs are increasingly used for imputation, anomaly detection, and synthetic data generation. In addition to extending a four-level causal intervention framework to four tabular and one image benchmark across five different VAE architectures (with 75 individual training runs per architecture and three random seed values for each run), this paper introduces three new techniques: posterior-calibration of Causal Effect Strength (CES), path-specific activation patching, and Feature-Group Disentanglement (FGD). The results from our experiments demonstrate that: (i) Tabular VAEs have circuits with modularity that is approximately 50% lower than their image counterparts. (ii) $β$-VAE experiences nearly complete collapse in CES scores when applied to heterogeneous tabular features (0.043 CES score for tabular data compared to 0.133 CES score for images), which can be directly attributed to reconstruction quality degradation (r = -0.886 correlation coefficient between CES and MSE). (iii) CES successfully captures nine of eleven statistically significant architecture differences using Holm--Šidák corrections. (iv) Interventions with high specificity predict the highest downstream AUC values (r = 0.460, p < .001). This study challenges the common assumption that architectural guidance from image-related studies can be transferred to tabular datasets.

可解释性表格数据变分自编码器因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。