用立体声实现声音事件定位,区分画面内外源。
Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification
- 基于立体声与视频输入,联合估计声源方位与距离。
- 引入画面内外分类任务,提升实际场景适用性。
- 新数据集与评估指标,适配普通音频媒体应用。
本文介绍了 DCASE2025 挑战赛第3项任务——立体声声音事件定位与检测(stereo SELD)的客观目标、数据集、基线系统及评估指标。与以往使用四通道一阶全向声学(FOA)或麦克风阵列不同,本年度挑战聚焦于更常见的立体声数据,将重点从全景声场分析转向有限视场(FOV)的日常音视频场景。由于立体声固有的角度模糊性,任务专注于水平方向(左右轴)到达角(DOA)估计与距离估计。挑战仍分为纯音频与音视频双赛道,音视频赛道新增了画面内外事件分类子任务,以应对视场限制。本文发布 DCASE2025 Task3 Stereo SELD Dataset,其立体声与视角视频片段源自 STARSS23 数据集的采样与转换。基线系统以立体声和对应视频帧为输入,除常规的声音事件分类与定位外,还集成画面内外分类功能。评估指标新增画面内外准确率,用于衡量模型对声源是否在画面内的判断能力。实验表明,基线系统在立体声数据上表现良好。
原文摘要 · Abstract (English)
This paper presents the objective, dataset, baseline, and metrics of Task 3 of the DCASE2025 Challenge on sound event localization and detection (SELD). In previous editions, the challenge used four-channel audio formats of first-order Ambisonics (FOA) and microphone array. In contrast, this year's challenge investigates SELD with stereo audio data (termed stereo SELD). This change shifts the focus from more specialized 360° audio and audiovisual scene analysis to more commonplace audio and media scenarios with limited field-of-view (FOV). Due to inherent angular ambiguities in stereo audio data, the task focuses on direction-of-arrival (DOA) estimation in the azimuth plane (left-right axis) along with distance estimation. The challenge remains divided into two tracks: audio-only and audiovisual, with the audiovisual track introducing a new sub-task of onscreen/offscreen event classification necessitated by the limited FOV. This challenge introduces the DCASE2025 Task3 Stereo SELD Dataset, whose stereo audio and perspective video clips are sampled and converted from the STARSS23 recordings. The baseline system is designed to process stereo audio and corresponding video frames as inputs. In addition to the typical SELD event classification and localization, it integrates onscreen/offscreen classification for the audiovisual track. The evaluation metrics have been modified to introduce an onscreen/offscreen accuracy metric, which assesses the models' ability to identify which sound sources are onscreen. In the experimental evaluation, the baseline system performs reasonably well with the stereo audio data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。