arXiv:2605.14736cs.SDcs.LG2026-05

用视觉和空间信息提升小麦克风阵列的语音提取效果

IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments

论文配图:IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments
图 1 · 摘自论文原文
  • 融合视觉、空间特征与深度网络,实现可选目标语音分离
  • 在-1至10 dB信噪比下达到9.31 dB SI-SDR,优于传统方法4.85 dB
  • 适合小型设备部署,尤其适用于复杂声学环境中的语音增强

目标语音提取在紧凑设备上仍具挑战,因单声道神经模型缺乏空间信息,而传统波束成形器在厘米级麦克风阵列下分辨力下降。本文提出IsoNet,一种面向4麦克风阵列的可选目标语音提取系统。该系统结合多通道STFT特征、GCC-PHAT空间线索、人脸条件视觉嵌入及方向估计辅助监督,集成于U-Net掩码估计网络。在25,000个模拟VoxCeleb混响样本上训练三种渐进式难度的课程变体。在-1至10 dB SNR的硬测试集上,IsoNet-CL1实现9.31 dB SI-SDR,较原始混合信号提升4.85 dB,PESQ为2.13,STOI为0.84。理想延迟求和与MVDR波束成形器分别使信号恶化4.82 dB和6.08 dB SI-SDRi,表明所提多模态学习方法解决了传统空间滤波失效的场景。消融实验显示视觉条件、GCC-PHAT特征与扩展时延窗编码均带来稳定增益。结果确立了受控仿真下紧凑阵列、人脸选择语音提取的基准,并指明真实部署的主要障碍:相位重建、多干扰源混合与仿真到现实的迁移问题。

原文摘要 · Abstract (English)

Target speech extraction remains difficult for compact devices because monaural neural models lack spatial evidence and classical beamformers lose resolving power when the microphone aperture is only a few centimetres. We present IsoNet, a user-selectable audio-visual target speech extraction system for a compact 4-microphone array. IsoNet combines complex multi-channel STFT features, GCC-PHAT spatial cues, face-conditioned visual embeddings, and auxiliary direction-of-arrival supervision inside a U-Net mask estimation network. Three curriculum variants were trained on 25,000 simulated VoxCeleb mixtures with progressively difficult SNR regimes. On a hard test set spanning -1 to 10 dB SNR, IsoNet-CL1 achieves 9.31 dB SI-SDR, a 4.85 dB improvement over the mixture, with PESQ 2.13 and STOI 0.84. Oracle delay-and-sum and MVDR beamformers degrade the same mixtures by 4.82 dB and 6.08 dB SI-SDRi, respectively, showing that the proposed learned multimodal conditioning solves a regime where conventional spatial filtering is ineffective. Ablation studies show consistent gains from visual conditioning, GCC-PHAT features, and extended delay-bin encoding. The results establish a compact-array, face-selectable speech extraction baseline under controlled simulation and identify the remaining barriers to real deployment, especially phase reconstruction, multi-interferer mixtures, and simulation-to-real transfer.

语音分离多模态小阵列视觉融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。