通过信号分解结构增强变分自编码器,提升时频特征的解耦表征能力。
Variational decomposition autoencoding improves disentanglement of latent representations
- 引入分解先验,让编码器显式学习时间-频率子空间
- 在语音、发音障碍评估等3个数据集上解耦性优于现有VAE方法
- 适合需要可解释动态信号表征的医学与人机交互场景
理解复杂、非平稳、高维随时间演化的信号结构是科学数据分析的核心挑战。在语音和生物医学信号处理等领域,学习解耦且可解释的表示对揭示潜在生成机制至关重要。传统无监督表示学习方法(如变分自编码器,VAEs)常难以捕捉此类数据中固有的时序与谱多样性。本文提出变分分解自编码(VDA),通过引入强信号分解结构先验扩展VAEs。VDA具体实现为分解变分自编码器(DecVAEs)——一种仅含编码器的神经网络,结合信号分解模型、对比自监督任务和变分先验近似,学习与时间-频率特性对齐的多个潜在子空间。我们在模拟数据及三个公开科学数据集(语音识别、发音障碍严重度评估、情感语音分类)上验证了DecVAEs的有效性。结果表明,DecVAEs在解耦质量、跨任务泛化能力和潜在编码可解释性方面均超越当前最优的基于VAE的方法。这些发现表明,具备分解感知能力的架构可作为从动态信号中提取结构化表示的稳健工具,具有临床诊断、人机交互和自适应神经技术等应用潜力。
原文摘要 · Abstract (English)
Understanding the structure of complex, nonstationary, high-dimensional time-evolving signals is a central challenge in scientific data analysis. In many domains, such as speech and biomedical signal processing, the ability to learn disentangled and interpretable representations is critical for uncovering latent generative mechanisms. Traditional approaches to unsupervised representation learning, including variational autoencoders (VAEs), often struggle to capture the temporal and spectral diversity inherent in such data. Here we introduce variational decomposition autoencoding (VDA), a framework that extends VAEs by incorporating a strong structural bias toward signal decomposition. VDA is instantiated through variational decomposition autoencoders (DecVAEs), i.e., encoder-only neural networks that combine a signal decomposition model, a contrastive self-supervised task, and variational prior approximation to learn multiple latent subspaces aligned with time-frequency characteristics. We demonstrate the effectiveness of DecVAEs on simulated data and three publicly available scientific datasets, spanning speech recognition, dysarthria severity evaluation, and emotional speech classification. Our results demonstrate that DecVAEs surpass state-of-the-art VAE-based methods in terms of disentanglement quality, generalization across tasks, and the interpretability of latent encodings. These findings suggest that decomposition-aware architectures can serve as robust tools for extracting structured representations from dynamic signals, with potential applications in clinical diagnostics, human-computer interaction, and adaptive neurotechnologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。