arXiv:2510.02500cs.SD2025-10中稿 · DCASE 2025 Worksho…

将对比学习融入生成模型,提升环境声的鲁棒表示能力

Latent Multi-view Learning for Robust Environmental Sound Representations

  • 用多视图框架分离声源与设备信息,通过对比学习引导信息流动
  • 在城市声音传感器数据集上,分类任务性能优于传统自监督方法
  • 可解耦环境声音属性,适合需要可解释表示的研究场景

自监督学习(SSL)方法如对比和生成式方法,已利用无标签数据推进环境声音表示学习。然而,这些方法如何在统一框架中互补仍研究不足。本文提出一种多视图学习框架,将对比原则融入生成流程,以捕捉声音源和设备信息。方法将压缩音频隐变量编码至视图特定与共用子空间,受两个自监督目标指导:对比学习用于子空间间定向信息流,重建损失用于整体信息保留。我们在城市声音传感器网络数据集上评估该方法,在声音源与传感器分类任务中表现优于传统SSL技术。此外,我们研究了模型在不同训练配置下,于结构化隐空间中解耦环境声音属性的潜力。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) approaches, such as contrastive and generative methods, have advanced environmental sound representation learning using unlabeled data. However, how these approaches can complement each other within a unified framework remains relatively underexplored. In this work, we propose a multi-view learning framework that integrates contrastive principles into a generative pipeline to capture sound source and device information. Our method encodes compressed audio latents into view-specific and view-common subspaces, guided by two self-supervised objectives: contrastive learning for targeted information flow between subspaces, and reconstruction for overall information preservation. We evaluate our method on an urban sound sensor network dataset for sound source and sensor classification, demonstrating improved downstream performance over traditional SSL techniques. Additionally, we investigate the model's potential to disentangle environmental sound attributes within the structured latent space under varied training configurations.

自监督学习声音表示多视图学习隐空间解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。