arXiv:2511.11825cs.SDcs.AI2025-11被引 1

用双输入融合视觉机制,实时提升复杂噪声中的语音清晰度

Real-Time Speech Enhancement via a Hybrid ViT: A Dual-Input Acoustic-Image Feature Fusion

  • 设计双输入结构,融合声学与图像特征建模时频依赖
  • 在多个数据集上显著提升语音可懂度与感知质量
  • 轻量化设计适合嵌入式设备部署,适用于真实场景

嘈杂环境中语音质量和可懂度严重下降。本文提出一种基于Transformer的新型学习框架,用于解决单通道噪声抑制的实时应用问题。尽管现有深度学习网络在处理平稳噪声方面表现优异,但在非平稳噪声(如狗吠、婴儿哭声)的真实环境中的性能常显著下降。所提出的混合视觉变压器(Hybrid ViT)框架通过双输入声学-图像特征融合,有效建模噪声信号中的时序与频谱依赖关系。该框架专为真实音频环境设计,计算量轻,适合嵌入式设备部署。采用四个标准指标(PESQ、STOI、Seg SNR、LLR)评估性能。实验基于Librispeech作为干净语音源,UrbanSound8K和Google Audioset作为噪声源,结果表明,相比原始噪声信号,该方法在降噪、语音可懂度和感知质量方面均有显著提升,性能接近纯净参考信号。

原文摘要 · Abstract (English)

Speech quality and intelligibility are significantly degraded in noisy environments. This paper presents a novel transformer-based learning framework to address the single-channel noise suppression problem for real-time applications. Although existing deep learning networks have shown remarkable improvements in handling stationary noise, their performance often diminishes in real-world environments characterized by non-stationary noise (e.g., dog barking, baby crying). The proposed dual-input acoustic-image feature fusion using a hybrid ViT framework effectively models both temporal and spectral dependencies in noisy signals. Designed for real-world audio environments, the proposed framework is computationally lightweight and suitable for implementation on embedded devices. To evaluate its effectiveness, four standard and commonly used quality measurements, namely PESQ, STOI, Seg SNR, and LLR, are utilized. Experimental results obtained using the Librispeech dataset as the clean speech source and the UrbanSound8K and Google Audioset datasets as the noise sources, demonstrate that the proposed method significantly improves noise reduction, speech intelligibility, and perceptual quality compared to the noisy input signal, achieving performance close to the clean reference.

语音增强Transformer实时系统双输入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。