arXiv:2508.06347cs.LGcs.AI2025-08被引 2

让表格数据的潜在表示更可解释,靠结构设计而非调参

Structural Equation-VAE: Disentangled Latent Representations for Tabular Data

  • 将测量结构嵌入VAE架构,按指标分组设计潜在空间
  • 在模拟数据上显著提升因子还原度和对干扰的鲁棒性
  • 适合科学社科中理论驱动、需验证测量有效性的场景

从表格数据学习可解释的潜在表示仍是深度生成建模中的挑战。本文提出SE-VAE(结构方程-变分自编码器),一种将测量结构直接融入变分自编码器设计的新架构。受结构方程模型启发,SE-VAE使潜在子空间与已知指标分组对齐,并引入全局噪声潜在变量以分离特定构念的混淆变异。该模块化架构通过设计实现解耦,而非仅依赖统计正则化。我们在一系列模拟表格数据集上评估了SE-VAE,使用标准解耦度量与多个领先基线对比。结果表明,SE-VAE在因子恢复、可解释性和对干扰变异的鲁棒性方面持续优于其他方法。消融实验显示,架构结构而非正则化强度是性能的关键驱动因素。SE-VAE为科学与社会领域中理论驱动的白盒生成建模提供了原则性框架,其中潜构念需有理论依据且测量有效性至关重要。

原文摘要 · Abstract (English)

Learning interpretable latent representations from tabular data remains a challenge in deep generative modeling. We introduce SE-VAE (Structural Equation-Variational Autoencoder), a novel architecture that embeds measurement structure directly into the design of a variational autoencoder. Inspired by structural equation modeling, SE-VAE aligns latent subspaces with known indicator groupings and introduces a global nuisance latent to isolate construct-specific confounding variation. This modular architecture enables disentanglement through design rather than through statistical regularizers alone. We evaluate SE-VAE on a suite of simulated tabular datasets and benchmark its performance against a series of leading baselines using standard disentanglement metrics. SE-VAE consistently outperforms alternatives in factor recovery, interpretability, and robustness to nuisance variation. Ablation results reveal that architectural structure, rather than regularization strength, is the key driver of performance. SE-VAE offers a principled framework for white-box generative modeling in scientific and social domains where latent constructs are theory-driven and measurement validity is essential.

表格式数据解耦表示可解释性结构方程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。