arXiv:2605.06108eess.AS2026-05被引 1

NDF+联合处理声源定向与混响分离,实现可调控的立体声差异。

NDF+: Joint Neural Directional Filtering and Diffuse Sound Extraction

  • 将定向麦克风重建分为去混响与漫射声提取两个耦合任务
  • 在混响环境下优于传统基线方法,且保持原始模型音质
  • 支持通过调节漫射成分控制左右声道音量差,适合空间音频应用

最近提出的神经定向滤波(NDF)是一种灵活的方法,用于重建具有特定指向性的虚拟定向麦克风(VDM),以实现空间声音捕获。在此基础上,本文提出NDF+,能够联合进行神经定向滤波与漫射声提取。NDF+将VDM估计重构为两个耦合子任务:去混响的VDM重建与漫射声提取。这一重构使NDF+可在最终重建的VDM输出中操控漫射成分。我们在混响条件下评估了NDF+,并与代表性传统基线进行了比较。结果表明,NDF+在两个子任务上均持续优于基线,同时保持与原单任务NDF模型相当的VDM重建质量。这些发现表明,NDF+在VDM重建中引入了额外的自由度以控制漫射声。在立体声录制应用中,NDF+可通过调整估计的漫射成分,实现可控的左右声道间电平差。

原文摘要 · Abstract (English)

Recently, neural directional filtering (NDF) has been introduced as a flexible approach for reconstructing a virtual directional microphone (VDM) with a desired directivity pattern for spatial sound capture. Building on this idea, we propose NDF+, which enables joint neural directional filtering and diffuse sound extraction. NDF+ reformulates VDM estimation into two coupled subtasks: dereverberated VDM reconstruction and diffuse sound extraction. This reformulation enables NDF+ to manipulate diffuse components in the final reconstructed VDM output. We evaluated NDF+ under reverberant conditions and compared it with representative conventional baselines. Results show that NDF+ consistently outperforms the baselines on both subtasks, while maintaining VDM reconstruction quality comparable to that of the original single-task NDF model. These findings indicate that NDF+ introduces an additional degree of freedom for diffuse sound control in the VDM reconstruction. In a stereo recording application, NDF+ provides controllable inter-channel level differences between left and right channels by adjusting the estimated diffuse component.

空间音频语音增强深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。