arXiv:2409.09619cs.SDeess.AS2024-09中稿 · ICASSP 2025被引 1

让声音模型学会分离不同声源,提升听觉理解的可解释性。

Compositional Audio Representation Learning

  • 设计了两种新方法:监督分类与无监督特征重建,实现声源独立表征
  • 监督学习比无监督更有效,特征重建优于频谱重建
  • 适合需要声音成分解析的场景,如多音轨分析、智能音频处理

人类听觉具有组合性——能从包含多个声源的复杂声音场景中识别出各个声音流。然而,现有方法通常使用片段级表示,无法分离组成声源。本文提出源中心音频表征学习方法,使每个声源由独立的嵌入向量表示。设计了两种新模型:一种基于分类任务的监督模型,另一种基于特征重建的无监督模型,两者均优于基线。通过音频分类任务评估设计选择,发现监督信号有助于学习源中心表征,且重建音频特征比重建频谱更有效。该方法可提升机器听觉的可解释性与解码灵活性。

原文摘要 · Abstract (English)

Human auditory perception is compositional in nature -- we identify auditory streams from auditory scenes with multiple sound events. However, such auditory scenes are typically represented using clip-level representations that do not disentangle the constituent sound sources. In this work, we learn source-centric audio representations where each sound source is represented using a distinct, disentangled source embedding in the audio representation. We propose two novel approaches to learning source-centric audio representations: a supervised model guided by classification and an unsupervised model guided by feature reconstruction, both of which outperform the baselines. We thoroughly evaluate the design choices of both approaches using an audio classification task. We find that supervision is beneficial to learn source-centric representations, and that reconstructing audio features is more useful than reconstructing spectrograms to learn unsupervised source-centric representations. Leveraging source-centric models can help unlock the potential of greater interpretability and more flexible decoding in machine listening.

音频表征源分离深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。