arXiv:2603.11242stat.MLcs.LG2026-03

统一多种生成模型,实现可解释性强的潜在空间解耦。

A Unified Latent Space Disentanglement VAE Framework with Robust Disentanglement Effectiveness Evaluation

  • 构建bfVAE框架,整合主流解耦VAE方法。
  • 提出FVH-LT与DBSR-LS,有效揭示语义相关潜在结构。
  • 设计LSSI指标,无需真实生成因子即可量化解耦效果。

评估和解释变分自编码器(VAEs)等模型的潜在表示仍是挑战,尤其在缺乏真实生成因子时。为此,本文将多种先进解耦VAE方法统一至一个框架——bfVAE。为评估解耦效果并提升潜在空间可解释性,提出特征方差异质性通过潜在遍历(FVH-LT)与潜在空间脏块稀疏回归(DBSR-LS)。为确保学习到的潜在空间可解释性稳定,开发贪心对齐策略(GAS),缓解标签交换问题,并实现跨运行潜在维度对齐,奠定结果聚合基础。进一步提出基于GAS对齐输出的标量潜在空间分离指数(LSSI),在无真实生成因子前提下总结整体潜在结构分离程度。在七个表格与图像数据集上,将bfVAE与五种VAE模型对比,验证了FVH-LT、DBSR-LS及LSSI的有效性。实验表明,bfVAE在解耦与重构之间取得更优综合平衡;FVH-LT与DBSR-LS能可靠揭示语义有意义且领域相关的潜在结构,结果一致;LSSI可有效定量总结潜在结构分离程度。

原文摘要 · Abstract (English)

Evaluating and interpreting latent representations, such as variational autoencoders (VAEs), remains a significant challenge for diverse data types, especially when ground-truth generative factors are unknown. To address this, we unify several state-of-the-art disentangled VAE approaches for latent space disentanglement into one framework -- bfVAE. To assess the effectiveness of a disentangled VAE model and enhance latent space interpretability, we propose Feature Variance Heterogeneity via Latent Traversal (FVH-LT) and Dirty Block Sparse Regression in Latent Space (DBSR-LS). To ensure robust interpretability of learned latent space, we develop a greedy alignment strategy (GAS) that mitigates label switching and aligns latent dimensions across runs to set the foundation of result aggregation. We also introduce a convenient scalar latent space separation index (LSSI) based on the GAS-aligned outputs of FVH-LT and DBSR-LS to summarize the overall latent structural separation without knowledge of the ground-truth generative factors. We compare bfVAE to five VAE models and validate the effectiveness FVH-LT, DBSR-LS, and LSSI in on seven tabular and image datasets. Under our examined experimental settings, bfVAE provides a more flexible disentanglement framework achieves more favorable overall trade-off between disentanglement and reconstruction than the benchmark VAE models; FVH-LT and DBSR-LS reliably uncover semantically meaningful and domain-relevant latent structures and generally yield consistent results; and LSSI makes an effective quantitative summary of latent structural separation.

解耦表征VAE可解释性潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。