arXiv:2509.14785cs.SD2025-09被引 6

让音频文本嵌入模型学会分辨多个声源的位置,提升复杂环境下的理解能力。

Spatial-CLAP: Learning Spatially-Aware audio--text Embeddings for Multi-Source Conditions

  • 引入内容感知的空间编码器,将声音内容与位置信息联合建模。
  • 提出空间对比学习策略,使模型在多声源下仍能准确对应音源与位置。
  • 首次在三声源未见混合场景中验证多源训练优势,适合音频定位与跨模态检索任务。

对比语言-音频预训练(CLAP)在音视频嵌入框架中取得了显著成果,但现有方法仅适用于单声道或单声源条件,难以充分捕捉空间信息。在多声源条件下,正确建立每个声源与其位置的对应关系是核心挑战。为此,我们提出Spatial-CLAP,引入一种内容感知的空间编码器,实现音频内容与空间表征的耦合。同时,提出空间对比学习(SCL),通过显式约束学习正确的音源-位置对应关系,提升多声源条件下的嵌入可靠性。实验评估涵盖下游任务,结果表明Spatial-CLAP在多声源条件下仍能学习有效嵌入,并验证了SCL的有效性。此外,在未见三声源混合场景中的评估突显了传统单声源训练与本研究所提多声源训练范式的根本差异。这些发现确立了空间感知音频-文本嵌入的新范式。

原文摘要 · Abstract (English)

Contrastive language--audio pretraining (CLAP) has achieved remarkable success as an audio--text embedding framework, but existing approaches are limited to monaural or single-source conditions and cannot fully capture spatial information. The central challenge in modeling spatial information lies in multi-source conditions, where the correct correspondence between each sound source and its location is required. To tackle this problem, we propose Spatial-CLAP, which introduces a content-aware spatial encoder that enables spatial representations coupled with audio content. We further propose spatial contrastive learning (SCL), a training strategy that explicitly enforces the learning of the correct correspondence and promotes more reliable embeddings under multi-source conditions. Experimental evaluations, including downstream tasks, demonstrate that Spatial-CLAP learns effective embeddings even under multi-source conditions, and confirm the effectiveness of SCL. Moreover, evaluation on unseen three-source mixtures highlights the fundamental distinction between conventional single-source training and our proposed multi-source training paradigm. These findings establish a new paradigm for spatially-aware audio--text embeddings.

音频定位跨模态空间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。