arXiv:2410.17457cs.SDeess.AS2024-10被引 5

用毫米波雷达窃听手机通话并自动转写,首次实现全语料大词表识别。

mmWave-Whisper: Phone Call Eavesdropping and Transcription Using Millimeter-Wave Radar

  • 通过雷达捕捉手机耳塞振动,转为音频再用Whisper模型识别
  • 在25至125厘米距离内实现44.74%词准确率、62.52%字符准确率
  • 突破以往仅限小词汇或扬声器的局限,警示雷达语音窃听新风险

本文提出mmWave-Whisper系统,展示利用现成调频连续波(FMCW)毫米波雷达远程窃听手机通话并实现全语料自动语音识别(ASR)的可行性。系统工作于77-81 GHz频段,通过捕捉智能手机耳塞的微振动,将其转化为音频信号,并经处理后生成语音转录。与以往仅针对扬声器或有限词汇的研究不同,这是首个针对手机耳塞振动实现大词汇量和完整句子识别的工作。系统通过合成数据训练、领域自适应及引入OpenAI Whisper模型,克服了缺乏大规模训练数据、信噪比低和雷达数据频率信息受限等挑战。在25至125厘米距离范围内,系统达到44.74%的词准确率和62.52%的字符准确率。论文揭示了人工智能技术快速发展带来的新型滥用风险。

原文摘要 · Abstract (English)

This paper introduces mmWave-Whisper, a system that demonstrates the feasibility of full-corpus automated speech recognition (ASR) on phone calls eavesdropped remotely using off-the-shelf frequency modulated continuous wave (FMCW) millimeter-wave radars. Operating in the 77-81 GHz range, mmWave-Whisper captures earpiece vibrations from smartphones, converts them into audio, and processes the audio to produce speech transcriptions automatically. Unlike previous work that focused on loudspeakers or limited vocabulary, this is the first work to perform such a speech recognition by handling large vocabulary and full sentences on earpiece vibrations from smartphones. This approach expands the potential of radar-audio eavesdropping. mmWave-Whisper addresses challenges such as the lack of large scale training datasets, low SNR, and limited frequency information in radar data through a systematic pipeline designed to leverage synthetic training data, domain adaptation, and inference by incorporating OpenAI's Whisper automatic speech recognition model. The system achieves a word accuracy rate of 44.74% and a character accuracy rate of 62.52% over a range of 25 cm to 125 cm. The paper highlights emerging misuse modalities of AI as the technology evolves rapidly.

雷达窃听语音识别安全威胁

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。