用跨模态嵌入提升视频中声音事件的立体定位精度。
Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
- 融合音频与视觉的对比语言对齐模型增强空间感知
- 在DCASE2025挑战赛中取得第二名,显著优于基线
- 适合关注多模态音频定位与预训练方法的研究者
本文针对常规视频内容中的三维声音事件定位与检测(3D SELD)任务,提出一种融合空间与语义嵌入的方法。该任务需同时完成时间上事件分类与空间定位,涉及跨空间、时间与语义维度的推理,其中语义建模尤为困难。传统SLED方法依赖多通道输入,受限于数据规模难以利用大规模预训练。为此,本文引入预训练的对比语言对齐模型:音频端使用CLAP,视觉端使用OWL-ViT,将其嵌入改进的Conformer模块,构建跨模态融合网络(Cross-Modal Conformer)。在DCASE2025 Task3 Stereo SELD数据集开发集上进行消融实验,评估各语言对齐模型贡献,并与基准系统对比。此外,详细说明了为预训练构建的大规模合成音视频数据集及其通过左右声道交换增强的流程。最终方法结合充分预训练、模型集成与视觉后处理,在DCASE2025挑战赛任务3(赛道B)中位列第二,验证了其有效性。未来将探索各模态贡献及结构优化。
原文摘要 · Abstract (English)
In this study, we address the multimodal task of stereo sound event localization and detection with source distance estimation (3D SELD) in regular video content. 3D SELD is a complex task that combines temporal event classification with spatial localization, requiring reasoning across spatial, temporal, and semantic dimensions. The last is arguably the most challenging to model. Traditional SELD approaches typically rely on multichannel input, limiting their capacity to benefit from large-scale pre-training due to data constraints. To overcome this, we enhance a standard SELD architecture with semantic information by integrating pre-trained, contrastive language-aligned models: CLAP for audio and OWL-ViT for visual inputs. These embeddings are incorporated into a modified Conformer module tailored for multimodal fusion, which we refer to as the Cross-Modal Conformer. We perform an ablation study on the development set of the DCASE2025 Task3 Stereo SELD Dataset to assess the individual contributions of the language-aligned models and benchmark against the DCASE Task 3 baseline systems. Additionally, we detail the curation process of large synthetic audio and audio-visual datasets used for model pre-training. These datasets were further expanded through left-right channel swapping augmentation. Our approach, combining extensive pre-training, model ensembling, and visual post-processing, achieved second rank in the DCASE 2025 Challenge Task 3 (Track B), underscoring the effectiveness of our method. Future work will explore the modality-specific contributions and architectural refinements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。