arXiv:2510.01940eess.AS2025-10被引 1

用变分自编码器无监督聚类声学环境,适合助听设备实时应用。

Clustering of Acoustic Environments with Variational Autoencoders for Hearing Devices

  • 基于变分自编码器与Gumbel-Softmax,实现声学环境的无监督聚类。
  • 在语音数字数据上聚类准确率高,在城市声景中仍保持有效性能。
  • 适用于内存受限的助听设备,支持时间窗口处理,部署友好。

传统声学环境分类依赖经典信号处理或有监督学习,前者难以提取高维数据的有效表示,后者受限于标签稀缺。由于人为标签未必反映真实声学场景结构,本文探索使用变分自编码器(VAE)进行无监督声学环境聚类。提出一种基于Gumbel-Softmax重参数化的类别型潜在变量聚类方法,结合时间上下文窗口策略以降低内存占用,适配真实助听设备场景。同时对音频聚类任务优化了VAE架构。在语音数字识别(标签有意义)和城市声景(时间频率重叠强)两个任务上验证:所有变分方法在语音数字上表现良好,唯独本模型在城市声景中实现有效聚类,得益于其类别型潜空间设计。

原文摘要 · Abstract (English)

Traditional acoustic environment classification relies on: i) classical signal processing algorithms, which are unable to extract meaningful representations of high-dimensional data; or on ii) supervised learning, limited by the availability of labels. Knowing that human-imposed labels do not always reflect the true structure of acoustic scenes, we explore the potential of (unsupervised) clustering of acoustic environments using variational autoencoders (VAEs). We employ a VAE model for categorical latent clustering with a Gumbel-Softmax reparameterization which can operate with a time-context windowing scheme for lower memory requirements, tailored for real-world hearing device scenarios. Additionally, general adaptations on VAE architectures for audio clustering are also proposed. The approaches are validated through the clustering of spoken digits, a simpler task where labels are meaningful, and urban soundscapes, where the recordings present strong overlap in time and frequency. While all variational methods succeeded when clustering spoken digits, only the proposed model achieved effective clustering performance on urban acoustic scenes, given its categorical nature.

声学聚类变分自编码器助听设备无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。