arXiv:2510.15669stat.MLcs.LG2025-10

用离散+连续潜变量分离多源信号,无需大量标注也能精准识别。

Disentanglement of Sources in a Multi-Stream Variational Autoencoder

  • 混合离散与连续潜变量,显式建模多源叠加机制
  • 仅用10%标签预训练即达高精度,复杂混合数字分离表现优异
  • 适用于手写数字叠加与双人对话语音分离,适合弱监督场景

变分自编码器(VAE)是学习解耦表征的主流方法,通常在单一连续潜空间中寻找解耦表示。本文提出并验证了一种新型多流变分自编码器(MS-VAE),通过结合离散与连续潜变量实现源分离。离散潜变量用于显式建模源组合,在解码器中叠加多个源。我们形式化定义了MS-VAE,推导其推理与学习方程,并数值验证其有效性。该模型高度灵活,可仅用少量监督(预训练使用部分标签后全无监督训练)。实验中,我们测试了其在叠加手写数字和声源分离中的能力:前者使用叠加MNIST数据集(常见基准),后者聚焦两人对话中的说话人区分任务。所有情况下均观察到清晰的源分离与竞争性性能;复杂混合数字(如三至四个数字)下表现尤为出色;说话人区分任务中漏检率低且说话人归属更精确。实验表明该方法在不同监督程度下均具灵活性,例如仅用10%标签预训练即可达到高性能。

原文摘要 · Abstract (English)

Variational autoencoders (VAEs) are among leading approaches to address the problem of learning disentangled representations. Typically a single VAE is used and disentangled representations are sought within its single continuous latent space. In this paper, we propose and provide a proof of concept for a novel Multi-Stream Variational Autoencoder (MS-VAE) that achieves disentanglement of sources by combining discrete and continuous latents. The discrete latents are used in an explicit source combination model, that superimposes a set of sources as part of the MS-VAE decoder. We formally define the MS-VAE approach, derive its inference and learning equations, and numerically investigate its principled functionality. The MS-VAE model is very flexible and can be trained using little supervision (we use fully unsupervised learning after pretraining with some labels). In our numerical experiments, we explored the ability of the MS-VAE approach in separating both superimposed hand-written digits as well as sound sources. For the former task we used superimposed MNIST digits (an increasingly common benchmark). For sound separation, our experiments focused on the task of speaker diarization in a recording conversation between two speakers. In all cases, we observe a clear separation of sources and competitive performance after training. For digit superpositions, performance is particularly competitive in complex mixtures (e.g., three and four digits). For the speaker diarization task, we observe an especially low rate of missed speakers and a more precise speaker attribution. Numerical experiments confirm the flexibility of the approach across varying amounts of supervision, and we observed high performance, e.g., when using just 10% of the labels for pretraining.

解耦表征多源分离弱监督变分自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。