用雷达感知生成对抗网络,从低信噪比毫米波信号中重建清晰语音。
mmWave Radar Aware Dual-Conditioned GAN for Speech Reconstruction of Signals With Low SNR
- 设计双条件生成模型,融合雷达特征与频谱信息增强输入
- 在-5dB至-1dB信噪比下实现优于现有方法的语音重建效果
- 适合对毫米波语音还原有需求的智能安防与远程通信场景
毫米波雷达采集信号带宽受限且噪声大,难以重建可懂的全频段语音。本文提出一种两阶段语音重建流程,基于雷达感知的双条件生成对抗网络(RAD-GAN),可在低信噪比(-5 dB至-1 dB)条件下、穿过玻璃墙的信号上实现带宽扩展。我们设计了针对毫米波特性的多梅尔判别器(MMD)和残差融合门(RFG),以增强生成器输入对多路条件信号的处理能力。两阶段流程包括:先在合成剪裁的干净语音上预训练,再在RFG生成的融合梅尔频谱上微调。实验表明,仅用有限数据集、无预训练模块、无数据增强的情况下,该方法仍优于当前最优方案。音频示例可访问 https://rad-gan-demo-site.vercel.app/。
原文摘要 · Abstract (English)
Millimeter-wave (mmWave) radar captures are band-limited and noisy, making for difficult reconstruction of intelligible full-bandwidth speech. In this work, we propose a two-stage speech reconstruction pipeline for mmWave using a Radar-Aware Dual-conditioned Generative Adversarial Network (RAD-GAN), which is capable of performing bandwidth extension on signals with low signal-to-noise ratios (-5 dB to -1 dB), captured through glass walls. We propose an mmWave-tailored Multi-Mel Discriminator (MMD) and a Residual Fusion Gate (RFG) to enhance the generator input to process multiple conditioning channels. The proposed two-stage pipeline involves pretraining the model on synthetically clipped clean speech and finetuning on fused mel spectrograms generated by the RFG. We empirically show that the proposed method, trained on a limited dataset, with no pre-trained modules, and no data augmentations, outperformed state-of-the-art approaches for this specific task. Audio examples of RAD-GAN are available online at https://rad-gan-demo-site.vercel.app/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。