用大模型预测上下文,让语音攻击更有效
Hearing the Unspoken: Language Model Priors for Acoustic Adversarial Attacks

- 用大语言模型实时预测语音上下文,突破时序限制
- 使语音识别错误率提升至35.6%,是之前最好结果的三倍
- 揭示了低延迟大模型如何被用来攻击实时语音系统
实时语音识别(ASR)系统在严格的时间约束下处理声学输入,决策基于不完整信息,这种因果性限制了攻击者的性能。本文提出的语义博弈攻击通过实时引入大语言模型生成的预测上下文,突破了这一因果限制。实验表明,该方法可使整体词错误率(Word Error Rate)达到35.6%,相较当前最先进水平提升三倍。该工作揭示了常见低延迟大语言模型工具如何被系统性地用于破坏实时语音识别流程。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) systems operating in real-time settings must process acoustic input under strict temporal constraints, where transcription decisions are inherently made on incomplete information. This causal constraint serves as an information bottleneck on attackers, significantly limiting attack performance. Our new Semantic Gambit attack breaks this causal limitation by augmenting the adversary with predictive context derived from a Large Language Model in real-time. Our experiments show that this form of augmentation can elevate the corpus-level Word Error Rate to 35.6% -- a three-fold increase over the current state-of-the-art. Ultimately, this work reveals how common, low-latency LLM tooling can be exploited to systematically subvert real-time ASR pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。