arXiv:2505.14476cs.CVcs.LG2025-05

让同一类样本的潜在表示共享活跃维度,提升可解释性。

Enhancing Interpretability of Sparse Latent Representations with Class Information

  • 引入新损失函数,强制同类别样本共享活跃的潜在维度。
  • 实现全局与类别特异性因素的共同捕捉,增强语义结构。
  • 适合需要可解释潜在空间的生成模型研究者使用。

变分自编码器(VAEs)是强大的生成模型,用于学习潜在表示。标准VAE通过使用所有维度生成分散且无结构的潜在空间,尤其在高维空间中限制了可解释性。为解决此问题,变分稀疏编码(VSC)引入了刺-板先验分布,使每个输入的潜在表示稀疏化,仅少数维度激活,从而更易解释。然而,VSC未能在同类别样本间建立一致的激活模式。本论文提出新方法,通过引入新损失函数,促使同类别样本共享相似的活跃维度,形成更具结构性的潜在空间。每个共享维度对应一个高层概念('因子'),既包含跨类全局因子,也捕捉类别特异性因子,显著提升潜在表示的实用性与可解释性。

原文摘要 · Abstract (English)

Variational Autoencoders (VAEs) are powerful generative models for learning latent representations. Standard VAEs generate dispersed and unstructured latent spaces by utilizing all dimensions, which limits their interpretability, especially in high-dimensional spaces. To address this challenge, Variational Sparse Coding (VSC) introduces a spike-and-slab prior distribution, resulting in sparse latent representations for each input. These sparse representations, characterized by a limited number of active dimensions, are inherently more interpretable. Despite this advantage, VSC falls short in providing structured interpretations across samples within the same class. Intuitively, samples from the same class are expected to share similar attributes while allowing for variations in those attributes. This expectation should manifest as consistent patterns of active dimensions in their latent representations, but VSC does not enforce such consistency. In this paper, we propose a novel approach to enhance the latent space interpretability by ensuring that the active dimensions in the latent space are consistent across samples within the same class. To achieve this, we introduce a new loss function that encourages samples from the same class to share similar active dimensions. This alignment creates a more structured and interpretable latent space, where each shared dimension corresponds to a high-level concept, or "factor." Unlike existing disentanglement-based methods that primarily focus on global factors shared across all classes, our method captures both global and class-specific factors, thereby enhancing the utility and interpretability of latent representations.

可解释性潜在空间生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。