arXiv:2506.00273eess.AScs.LG2025-06被引 2

用方向和语义信息从全景声中精准提取目标声音

SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction

  • 输入输出均为全景声,通过方向与语义联合引导
  • 在近邻干扰源场景下,语义+方向比仅方向更优
  • 适合沉浸式音频处理与智能音源分离任务

本文提出SoundSculpt,一种基于全景声输入输出的神经网络,用于从全景声混合信号中提取目标声场。该模型同时依赖空间信息(如指向沉浸式视频的目标方向)和语义嵌入(如由图像分割与描述生成的特征)。在合成与真实全景声混合数据上训练和评估,结果表明其性能优于多种信号处理基线方法。研究发现,仅使用空间条件已具有效性,但在存在空间邻近的次要声源时,结合空间与语义信息可显著提升分离效果。此外,本文对比了两种通过文本编码器从目标声音描述中生成的语义嵌入方式。

原文摘要 · Abstract (English)

This paper introduces SoundSculpt, a neural network designed to extract target sound fields from ambisonic recordings. SoundSculpt employs an ambisonic-in-ambisonic-out architecture and is conditioned on both spatial information (e.g., target direction obtained by pointing at an immersive video) and semantic embeddings (e.g., derived from image segmentation and captioning). Trained and evaluated on synthetic and real ambisonic mixtures, SoundSculpt demonstrates superior performance compared to various signal processing baselines. Our results further reveal that while spatial conditioning alone can be effective, the combination of spatial and semantic information is beneficial in scenarios where there are secondary sound sources spatially close to the target. Additionally, we compare two different semantic embeddings derived from a text description of the target sound using text encoders.

全景声音源分离语义引导空间感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。