用变分自编码器实现音频时频域无监督聚类,捕捉复杂声学模式。
Unsupervised Variational Acoustic Clustering
- 结合变分推断与高斯混合先验,构建时频域音频聚类模型
- 在语音数字数据集上性能显著优于传统方法
- 适合需要自动识别音频类别的研究者使用
我们提出一种用于时频域音频数据无监督聚类的变分声学聚类模型。该模型将变分推断扩展至自编码器框架,并以高斯混合模型作为潜在空间的先验。专为音频应用设计,引入了优化于时频处理的卷积-循环变分自编码器。实验结果表明,在语音数字数据集上,相比传统方法,该模型在准确率和聚类性能上均有显著提升,展现出对复杂音频模式更强的捕捉能力。
原文摘要 · Abstract (English)
We propose an unsupervised variational acoustic clustering model for clustering audio data in the time-frequency domain. The model leverages variational inference, extended to an autoencoder framework, with a Gaussian mixture model as a prior for the latent space. Specifically designed for audio applications, we introduce a convolutional-recurrent variational autoencoder optimized for efficient time-frequency processing. Our experimental results considering a spoken digits dataset demonstrate a significant improvement in accuracy and clustering performance compared to traditional methods, showcasing the model's enhanced ability to capture complex audio patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。