arXiv:2504.07652eess.AS2025-04被引 1

用类别分布强化音频聚类,解决重叠频谱的噪声场景聚类难题

Categorical Unsupervised Variational Acoustic Clustering

  • 引入类别分布与Gumbel-Softmax实现可微分聚类
  • 在城市声景数据上达到优异聚类性能,即使频谱高度重叠
  • 适合需要无监督音频分类的声学场景分析任务

我们提出一种基于类别分布的无监督变分音频聚类方法,用于时频域音频数据。通过引入类别分布,即使在时间与频率上存在强重叠的情况下,也能实现更清晰的聚类结果,这在多数城市声景数据中普遍存在。为此,我们采用Gumbel-Softmax分布作为类别分布的软近似,从而支持反向传播训练。在此框架中,softmax温度是调节聚类性能的主要机制。实验结果表明,该模型在所有测试数据集上均表现出色,即便在时间与频率高度重叠的条件下仍能取得优异聚类效果。

原文摘要 · Abstract (English)

We propose a categorical approach for unsupervised variational acoustic clustering of audio data in the time-frequency domain. The consideration of a categorical distribution enforces sharper clustering even when data points strongly overlap in time and frequency, which is the case for most datasets of urban acoustic scenes. To this end, we use a Gumbel-Softmax distribution as a soft approximation to the categorical distribution, allowing for training via backpropagation. In this settings, the softmax temperature serves as the main mechanism to tune clustering performance. The results show that the proposed model can obtain impressive clustering performance for all considered datasets, even when data points strongly overlap in time and frequency.

音频聚类变分方法无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。