融合语义与空间信息,提升视频中声音事件定位与检测效果
Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
- 用预训练模型融合音视频语义特征,增强多模态理解
- 在开发集上显著超越基线,左/右声道交换增强数据泛化能力
- 适合做音视频事件分析、智能监控等应用的开发者参考
本文介绍我们提交至DCASE2025任务3挑战赛音频仅和音视频双轨的立体声事件定位与检测(SELD)系统。SELD是结合时间分类与空间定位的复杂任务,需跨时空语义维度推理,其中语义建模最具挑战性。传统SELD架构依赖多通道输入,受限于数据规模难以利用大规模预训练。为此,我们通过引入预训练对比语言对齐模型——音频端使用CLAP,视觉端使用OWL-ViT,将语义嵌入融入改进的Conformer模块,构建跨模态融合网络(Cross-Modal Conformer)。同时,引入基于自相关性的声学特征以提升距离估计精度。模型在筛选后的合成音视频数据集上进行预训练,并采用左右声道交换增强策略扩充训练数据。音频仅与音视频双模系统在开发集上均显著优于基线表现,性能进一步通过模型集成及基于人体关键点的视觉后处理提升。未来工作将分析各模态贡献并探索架构变体。
原文摘要 · Abstract (English)
This report presents our systems submitted to the audio-only and audio-visual tracks of the DCASE2025 Task 3 Challenge: Stereo Sound Event Localization and Detection (SELD) in Regular Video Content. SELD is a complex task that combines temporal event classification with spatial localization, requiring reasoning across spatial, temporal, and semantic dimensions. The last is arguably the most challenging to model. Traditional SELD architectures rely on multichannel input, which limits their ability to leverage large-scale pre-training due to data constraints. To address this, we enhance standard SELD architectures with semantic information by integrating pre-trained, contrastive language-aligned models: CLAP for audio and OWL-ViT for visual inputs. These embeddings are incorporated into a modified Conformer module tailored for multimodal fusion, which we refer to as the Cross-Modal Conformer. Additionally, we incorporate autocorrelation-based acoustic features to improve distance estimation. We pre-train our models on curated synthetic audio and audio-visual datasets and apply a left-right channel swapping augmentation to further increase the training data. Both our audio-only and audio-visual systems substantially outperform the challenge baselines on the development set, demonstrating the effectiveness of our strategy. Performance is further improved through model ensembling and a visual post-processing step based on human keypoints. Future work will investigate the contribution of each modality and explore architectural variants to further enhance results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。