arXiv:2603.25776stat.MLcs.LG2026-03被引 1

用自适应隐马尔可夫先验实现无监督盲源分离,让模型自己学会区分不同声源。

SAHMM-VAE: A Source-Wise Adaptive Hidden Markov Prior Variational Autoencoder for Unsupervised Blind Source Separation

  • 为每个潜在维度设计动态切换的隐马尔可夫先验,适配不同声源的时间结构。
  • 在训练中直接学习分离,无需后处理,可恢复原始声源信号。
  • 适用于需要可解释性建模的音频分离任务,尤其适合多源信号场景。

我们提出SAHMM-VAE,一种针对无监督盲源分离的源级自适应隐马尔可夫先验变分自编码器。不同于将潜在先验视为单一通用正则项,该框架为每个潜在维度分配其自身的自适应状态切换先验,使不同潜在维度在训练中被拉向不同的源特异性时序结构。在此设定下,源分离不作为外部后处理步骤,而是嵌入到变分学习过程本身。编码器、解码器、后验参数与源级先验参数联合优化,其中编码器逐步学习一个近似于混合变换逆映射的推理映射,而解码器扮演生成式混合模型角色。通过这种耦合优化,后验源轨迹与异构隐马尔可夫先验之间的渐进对齐成为不同潜在维度分离为不同源成分的机制。为实现这一思想,我们在同一框架内构建三个分支:高斯发射隐马尔可夫先验、马尔可夫切换自回归隐马尔可夫先验,以及带有状态自回归流变换的状态流转先验。实验表明,该框架在无监督条件下实现了源信号恢复,并学习到有意义的源级切换结构。更广泛地,该方法将结构化先验VAE从平滑、混合型及流型潜在先验拓展至自适应切换先验,为未来可解释且可能可识别的潜在源建模提供了基础。

原文摘要 · Abstract (English)

We propose SAHMM-VAE, a source-wise adaptive Hidden Markov prior variational autoencoder for unsupervised blind source separation. Instead of treating the latent prior as a single generic regularizer, the proposed framework assigns each latent dimension its own adaptive regime-switching prior, so that different latent dimensions are pulled toward different source-specific temporal organizations during training. Under this formulation, source separation is not implemented as an external post-processing step; it is embedded directly into variational learning itself. The encoder, decoder, posterior parameters, and source-wise prior parameters are optimized jointly, where the encoder progressively learns an inference map that behaves like an approximate inverse of the mixing transformation, while the decoder plays the role of the generative mixing model. Through this coupled optimization, the gradual alignment between posterior source trajectories and heterogeneous HMM priors becomes the mechanism through which different latent dimensions separate into different source components. To instantiate this idea, we develop three branches within one common framework: a Gaussian-emission HMM prior, a Markov-switching autoregressive HMM prior, and an HMM state-flow prior with state-wise autoregressive flow transformations. Experiments show that the proposed framework achieves unsupervised source recovery while also learning meaningful source-wise switching structures. More broadly, the method extends our structured-prior VAE line from smooth, mixture-based, and flow-based latent priors to adaptive switching priors, and provides a useful basis for future work on interpretable and potentially identifiable latent source modeling.

盲源分离隐马尔可夫变分自编码器无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。