arXiv:2504.02988cs.SDeess.AS2025-04被引 7

生成带空间对齐音效的360度合成视频,提升声音定位检测性能。

Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection

  • 用真实背景图合成360视频,确保音画空间一致。
  • 训练模型定位召回率达56.4%,误差21.9度。
  • 适合做声音定位与检测研究的团队使用。

我们提出SELDVisualSynth,一个用于音频-视觉声音事件定位与检测(SELD)的合成视频生成工具。该方法结合真实背景图像,提升合成音视频数据的真实性,同时保证音视频的空间对齐。工具生成360度合成视频,其中物体运动与合成的SELD音频数据及标注相匹配。实验表明,使用该数据训练的模型在多个指标上表现更优,定位召回率达到56.4%(LR),定位误差为21.9度(LE)。我们已开源该数据生成工具,以供SELD研究社区广泛使用。

原文摘要 · Abstract (English)

We present SELDVisualSynth, a tool for generating synthetic videos for audio-visual sound event localization and detection (SELD). Our approach incorporates real-world background images to improve realism in synthetic audio-visual SELD data while also ensuring audio-visual spatial alignment. The tool creates 360 synthetic videos where objects move matching synthetic SELD audio data and its annotations. Experimental results demonstrate that a model trained with this data attains performance gains across multiple metrics, achieving superior localization recall (56.4 LR) and competitive localization error (21.9deg LE). We open-source our data generation tool for maximal use by members of the SELD research community.

声音定位360视频数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。