生成带空间对齐音效的360度合成视频,提升声音定位检测性能。
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
- 用真实背景图合成360视频,确保音画空间一致。
- 训练模型定位召回率达56.4%,误差21.9度。
- 适合做声音定位与检测研究的团队使用。
我们提出SELDVisualSynth,一个用于音频-视觉声音事件定位与检测(SELD)的合成视频生成工具。该方法结合真实背景图像,提升合成音视频数据的真实性,同时保证音视频的空间对齐。工具生成360度合成视频,其中物体运动与合成的SELD音频数据及标注相匹配。实验表明,使用该数据训练的模型在多个指标上表现更优,定位召回率达到56.4%(LR),定位误差为21.9度(LE)。我们已开源该数据生成工具,以供SELD研究社区广泛使用。
原文摘要 · Abstract (English)
We present SELDVisualSynth, a tool for generating synthetic videos for audio-visual sound event localization and detection (SELD). Our approach incorporates real-world background images to improve realism in synthetic audio-visual SELD data while also ensuring audio-visual spatial alignment. The tool creates 360 synthetic videos where objects move matching synthetic SELD audio data and its annotations. Experimental results demonstrate that a model trained with this data attains performance gains across multiple metrics, achieving superior localization recall (56.4 LR) and competitive localization error (21.9deg LE). We open-source our data generation tool for maximal use by members of the SELD research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。