通过空间相关建模提升复杂场景下的目标声音提取效果
SoundCompass: Navigating Target Sound Extraction With Effective Directional Clue Integration In Complex Acoustic Scenes
- 用SPIN模块捕捉多通道频谱域中的跨通道空间关联
- 融合球谐编码的到达方向信息,实现更精准的空间定位
- 适合需要高精度声源分离的智能音频系统开发者
目标声音提取(TSE)近年来利用到达方向(DoA)提供的空间线索,该线索在任何声学场景中均存在。然而,以往基于DoA的方法依赖手工特征或离散编码,丢失了细粒度空间信息并限制了适应性。本文提出SoundCompass,一个以谱对间交互(SPIN)模块为核心的定向线索融合框架,能够捕获多通道信号在复频谱域中的跨通道空间相关性,从而保留完整的空间信息。输入特征以空间相关性表示,与以球谐(SH)编码形式表达的DoA线索在重叠频率子带上进行融合,继承了先前带分割架构的优点。此外,引入迭代精炼策略链式推理(CoI),递归地将DoA与前一阶段估计的声音事件激活融合。实验表明,SoundCompass结合SPIN、SH嵌入和CoI,在多种信号类别和空间配置下均能鲁棒地提取目标声源。
原文摘要 · Abstract (English)
Recent advances in target sound extraction (TSE) utilize directional clues derived from direction of arrival (DoA), which represent an inherent spatial property of sound available in any acoustic scene. However, previous DoA-based methods rely on hand-crafted features or discrete encodings, which lose fine-grained spatial information and limit adaptability. We propose SoundCompass, an effective directional clue integration framework centered on a Spectral Pairwise INteraction (SPIN) module that captures cross-channel spatial correlations in the complex spectrogram domain to preserve full spatial information in multichannel signals. The input feature expressed in terms of spatial correlations is fused with a DoA clue represented as spherical harmonics (SH) encoding. The fusion is carried out across overlapping frequency subbands, inheriting the benefits reported in the previous band-split architectures. We also incorporate the iterative refinement strategy, chain-of-inference (CoI), in the TSE framework, which recursively fuses DoA with sound event activation estimated from the previous inference stage. Experiments demonstrate that SoundCompass, combining SPIN, SH embedding, and CoI, robustly extracts target sources across diverse signal classes and spatial configurations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。