arXiv:2506.11125cs.CLcs.AI2025-06

用自然回声干扰语音识别,防诈骗电话却让人听不累。

ASRJam: Human-Friendly AI Speech Jamming to Prevent Automated Phone Scams

  • 用回声、混响等自然失真扰乱自动语音识别
  • 39人实验显示干扰效果最佳且人耳可接受
  • 专为防诈骗设计,不影响真人通话体验

大型语言模型(LLMs)结合文本转语音(TTS)与自动语音识别(ASR),正被广泛用于自动化语音诈骗(vishing)。这些系统可扩展性强且极具欺骗性,构成严重安全威胁。我们识别出ASR转录环节是诈骗链中最脆弱的环节,提出ASRJam防御框架,通过向受害者音频注入对抗性扰动,破坏攻击者的ASR系统,从而切断诈骗反馈环路,同时不影响人类通话理解。现有对抗音频技术往往令人不适且难以实时应用,因此我们进一步提出EchoGuard,一种利用自然失真(如混响、回声)的新型干扰器,对ASR具有强破坏性但对人类听感可接受。通过39人用户研究,对比三种先进攻击方法,结果表明EchoGuard在整体效用上表现最优,兼顾了ASR干扰能力和人类听觉体验。

原文摘要 · Abstract (English)

Large Language Models (LLMs), combined with Text-to-Speech (TTS) and Automatic Speech Recognition (ASR), are increasingly used to automate voice phishing (vishing) scams. These systems are scalable and convincing, posing a significant security threat. We identify the ASR transcription step as the most vulnerable link in the scam pipeline and introduce ASRJam, a proactive defence framework that injects adversarial perturbations into the victim's audio to disrupt the attacker's ASR. This breaks the scam's feedback loop without affecting human callers, who can still understand the conversation. While prior adversarial audio techniques are often unpleasant and impractical for real-time use, we also propose EchoGuard, a novel jammer that leverages natural distortions, such as reverberation and echo, that are disruptive to ASR but tolerable to humans. To evaluate EchoGuard's effectiveness and usability, we conducted a 39-person user study comparing it with three state-of-the-art attacks. Results show that EchoGuard achieved the highest overall utility, offering the best combination of ASR disruption and human listening experience.

语音对抗反诈骗人机共存

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。