用语言描述从立体声中分离特定声音并定位方向
LuSeeL: Language-queried Binaural Universal Sound Event Extraction and Localization
- 通过立体声的空间线索实现文本驱动的声音分离
- 在AudioCaps上性能显著优于单通道基线
- 同时完成声音提取与方向估计,适合复杂场景应用
大多数通用声音提取算法专注于从单通道音频混合中分离目标声音事件。然而真实世界是三维的,立体声(binaural audio)能模拟人类听觉,捕捉更丰富的空间信息,包括声源位置。这种空间上下文对于理解与建模复杂听觉场景至关重要,因为它能自然地提供声音检测与提取的信息。本文提出一种语言驱动的通用声音提取网络(LuSeeL),通过有效利用立体声信号中的空间线索,从立体声混合中分离出由文本描述的声音事件。同时,该模型联合预测目标声音的到达方向(DoA),利用提取网络中的空间特征实现。这种双任务方法利用互补的位置信息提升提取性能,并实现精确的DoA估计。在真实环境音频数据集AudioCaps上的实验结果表明,所提的LuSeeL模型显著优于单通道和单任务基线。
原文摘要 · Abstract (English)
Most universal sound extraction algorithms focus on isolating a target sound event from single-channel audio mixtures. However, the real world is three-dimensional, and binaural audio, which mimics human hearing, can capture richer spatial information, including sound source location. This spatial context is crucial for understanding and modeling complex auditory scenes, as it inherently informs sound detection and extraction. In this work, we propose a language-driven universal sound extraction network that isolates text-described sound events from binaural mixtures by effectively leveraging the spatial cues present in binaural signals. Additionally, we jointly predict the direction of arrival (DoA) of the target sound using spatial features from the extraction network. This dual-task approach exploits complementary location information to improve extraction performance while enabling accurate DoA estimation. Experimental results on the in-the-wild AudioCaps dataset show that our proposed LuSeeL model significantly outperforms single-channel and uni-task baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。