arXiv:2605.21141eess.AS2026-05

深度学习实现多说话人场景下的定向波束成形,提升语音增强效果。

Linearly Constrained Deep Beamformer for Multi-Speaker Scenarios

  • 用DNN直接从多通道噪声输入估计波束权重,通过动态加权损失函数满足空间约束。
  • 在相同空间特征下,性能优于经典LCMV波束成形器,且旁瓣更受控、背景噪声抑制更强。
  • 适合需要精准声源定位与干扰抑制的语音增强任务,如会议系统、智能音箱。

我们提出一种深度波束成形框架,用于在多说话人环境中增强目标说话人语音。一个深度神经网络(DNN)被训练为直接从含噪多通道输入估计波束权重,同时通过自适应多级损失函数逐步增加约束权重,以满足线性空间约束。该损失函数结合信号重建项和惩罚项,强制对目标方向保持无失真响应,并抑制干扰子空间。模型还利用目标相对传递函数(RTF)和估计的干扰子空间进行引导。所提模型可在指向目标说话人同时对干扰源形成零点,相比使用相同估计空间特征构造的经典LCMV波束成形器,整体增强性能更优。此外,与LCMV波束成形器相比,该模型具有更可控的旁瓣和更强的背景噪声抑制能力。

原文摘要 · Abstract (English)

We propose a deep beamforming framework for enhancing target speaker(s) in multi-speaker environments. A deep neural network (DNN) is trained to estimate beamforming weights directly from noisy multichannel inputs while satisfying linear spatial constraints through an adaptive multi-term loss with progressively increasing constraint weights. The loss combines signal reconstruction with penalties that enforce a distortionless response toward the target and suppress the interference subspace. The model is further guided by the target relative transfer function (RTF) and the estimated interference subspace. The proposed model can direct a beam toward the target speaker while directing nulls toward the interfering sources, achieving superior overall enhancement performance compared with the classical LCMV beamformer constructed by the same estimated spatial signatures. Furthermore, compared with the LCMV beamformer, the proposed model produces more controlled sidelobes and improved background-noise attenuation.

波束成形语音增强深度学习多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。