arXiv:2511.06205cs.SD2025-11被引 5

用毫米波雷达窃听扬声器振动,隔墙复现清晰语音

We Can Hear You with mmWave Radar! An End-to-End Eavesdropping System

  • 利用毫米波雷达捕捉扬声器振动信号,无需视线或侵入设备
  • 在无先验信息下重建可理解语音,音质优于现有方法
  • 适配真实场景,对未知说话人和复杂环境泛化能力强

随着语音交互技术普及,扬声器播放带来日益严重的语音隐私风险。传统窃听方法常需侵入或视线条件,实用性受限。本文提出mmSpeech,一种基于毫米波雷达的端到端窃听系统,仅凭扬声器播放引发的振动信号即可重建清晰语音,即使隔着墙壁且无说话人先验信息。我们揭示了最佳振动材料与雷达采样率组合,以高保真捕捉窄带毫米波信号。设计深度神经网络从估计的噪声频谱图中重构语音。为支持下游语音理解,引入合成训练流程,并微调预训练语音识别模型的编码器。使用商用毫米波雷达实现系统并进行大量实验验证。结果表明,mmSpeech在语音质量上达到当前最优水平,且对未见说话人及多种场景具有良好泛化能力。

原文摘要 · Abstract (English)

With the rise of voice-enabled technologies, loudspeaker playback has become widespread, posing increasing risks to speech privacy. Traditional eavesdropping methods often require invasive access or line-of-sight, limiting their practicality. In this paper, we present mmSpeech, an end-to-end mmWave-based eavesdropping system that reconstructs intelligible speech solely from vibration signals induced by loudspeaker playback, even through walls and without prior knowledge of the speaker. To achieve this, we reveal an optimal combination of vibrating material and radar sampling rate for capturing high-quality vibrations using narrowband mmWave signals. We then design a deep neural network that reconstructs intelligible speech from the estimated noisy spectrograms. To further support downstream speech understanding, we introduce a synthetic training pipeline and selectively fine-tune the encoder of a pre-trained ASR model. We implement mmSpeech with a commercial mmWave radar and validate its performance through extensive experiments. Results show that mmSpeech achieves state-of-the-art speech quality and generalizes well across unseen speakers and various conditions.

毫米波雷达语音安全隐私泄露非接触感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。