毫米波雷达可远程分辨多人对话中谁在说话,精度高达99%。
Who Speaks What from Afar: Eavesdropping In-Person Conversations via mmWave Sensing
- 利用物体振动差异,无须事先信息即可区分不同说话人。
- 在多参与场景下,语音分类准确率达0.99,信号增强效果稳定。
- 适合研究无线安全与隐私防护的学者,尤其关注物理层攻击者。
多人会议广泛存在于商业谈判、医疗咨询等领域,常涉及商业机密、战略计划及患者病情等敏感信息。已有研究表明,外部攻击者可通过毫米波雷达探测物体因语音振动产生的微小变化,实现会话窃听。然而,现有方法无法区分多人对话中各发言者的具体语义内容,易导致误判与决策失误。本文首次解决“谁在说什么”这一核心问题。通过利用环境中常见物体带来的空间多样性,提出一种无需预先知晓参与者身份、人数或座位布局的远程窃听系统。由于参会者位置不同,其语音会在邻近物体上引发独特的振动模式。为此,设计了一种抗噪的无监督频域分析方法,以区分不同说话人;同时引入深度学习框架,融合多个物体信号实现语音质量增强。通过大量实验验证了该攻击在语音分类与信号增强方面的可行性。结果表明,在多人会议环境中,语音分类准确率最高可达0.99;且在不同雷达-物体距离条件下,均保持稳定的语音增强效果。
原文摘要 · Abstract (English)
Multi-participant meetings occur across various domains, such as business negotiations and medical consultations, during which sensitive information like trade secrets, business strategies, and patient conditions is often discussed. Previous research has demonstrated that attackers with mmWave radars outside the room can overhear meeting content by detecting minute speech-induced vibrations on objects. However, these eavesdropping attacks cannot differentiate which speech content comes from which person in a multi-participant meeting, leading to potential misunderstandings and poor decision-making. In this paper, we answer the question ``who speaks what''. By leveraging the spatial diversity introduced by ubiquitous objects, we propose an attack system that enables attackers to remotely eavesdrop on in-person conversations without requiring prior knowledge, such as identities, the number of participants, or seating arrangements. Since participants in in-person meetings are typically seated at different locations, their speech induces distinct vibration patterns on nearby objects. To exploit this, we design a noise-robust unsupervised approach for distinguishing participants by detecting speech-induced vibration differences in the frequency domain. Meanwhile, a deep learning-based framework is explored to combine signals from objects for speech quality enhancement. We validate the proof-of-concept attack on speech classification and signal enhancement through extensive experiments. The experimental results show that our attack can achieve the speech classification accuracy of up to $0.99$ with several participants in a meeting room. Meanwhile, our attack demonstrates consistent speech quality enhancement across all real-world scenarios, including different distances between the radar and the objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。