arXiv:2601.17611eess.AScs.LG2026-01

用专家协作框架提升视频中声音事件的三维定位与距离估计

ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video

  • 分三个专业模块分别处理空间-语言、空间-时间、时间-语言关系
  • 在DCASE2025数据集上各项指标均优于现有方法
  • 适合做音视频多模态时空分析的研究者参考

视频中的声音事件三维定位与距离估计(3D SELD)任务要求在每一帧中识别出活跃的声音事件并估计其空间坐标。该多模态任务需在语义、空间和时间维度上进行联合推理,单一模型常难以有效应对。为此,我们提出专家团队(ToS)集成框架,包含三个互补的子网络:空间-语言模型、空间-时间模型和时间-语言模型。每个子网络专注于一对维度,贡献独特见解,如同具有不同专长的团队成员协同工作。ToS在DCASE2025任务3立体声SLED开发集上,相较于当前先进音视频模型,持续在关键指标上表现更优。未来工作将通过适配任务、训练策略及预训练方案强化各专家模块。

原文摘要 · Abstract (English)

Sound event localization and detection with distance estimation (3D SELD) in video involves identifying active sound events at each time frame while estimating their spatial coordinates. This multimodal task requires joint reasoning across semantic, spatial, and temporal dimensions, a challenge that single models often struggle to address effectively. To tackle this, we introduce the Team of Specialists (ToS) ensemble framework, which integrates three complementary sub-networks: a spatio-linguistic model, a spatio-temporal model, and a tempo-linguistic model. Each sub-network specializes in a unique pair of dimensions, contributing distinct insights to the final prediction, akin to a collaborative team with diverse expertise. ToS has been benchmarked against state-of-the-art audio-visual models for 3D SELD on the DCASE2025 Task 3 Stereo SELD development set, consistently outperforming existing methods across key metrics. Future work will extend this proof of concept by strengthening the specialists with appropriate tasks, training, and pre-training curricula.

三维定位音视频融合多模态分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。