实测超800万次语音攻击,发现物理声学环境极大提升攻击成功率。
Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks

- 构建高通量真实场景声学仿真框架,模拟物理世界攻击效果。
- 在Whisper和wav2vec上实现最高94.5%的词错误率提升。
- 提出双形式信噪比模型,分离隐蔽性与攻击有效性,适合安全研究者。
随着语音控制成为人机交互的常见方式,其面临的安全风险仍不清晰。这部分源于将纯数字对抗流程扩展到物理世界的困难。现有方法常忽略声学可探测性及几何影响,导致风险评估失真。本文通过真实测试、理论分析与新型高通量现实仿真框架,揭示这些问题。在超过800万次对抗评估中,我们证明声学意识可使Whisper和wav2vec的相对词错误率提升达94.5%。利用该框架,我们正式定义并实现了双形式信噪比(Dual-Form SNR),以解耦源端隐蔽性与目标攻击效能,克服当前研究的关键局限。该工作为可重复、可验证的声学安全研究提供了基础,强调应拥抱而非抽象声学环境。
原文摘要 · Abstract (English)
While voice control is rapidly becoming a ubiquitous vector of human-AI communication, the risks facing these systems remain poorly understood. This is, in part, a product of the difficulties in scaling strictly digital adversarial workflows to the physical world. These scale barriers have led the community to abstract away key acoustic factors relating to detectability and the influence of geometry on acoustics. These methodological and metrological shortcomings undermine our understanding of risk. We illuminate these issues through real-world testing, conceptual discussions, and a novel, high-throughput reality simulation framework. By testing over 8 million adversarial evaluations, we demonstrate that acoustic awareness yields relative Word Error Rate increases of up to 94.5\% under Whisper and wav2vec. We employ this framework to explore a formalize and operationalize a Dual-Form Signal to Noise Ratio to decouple source stealth from victim attack efficacy, resolving a crucial limitation in current works. This lays the groundwork for repeatable, verifiable research that embraces, rather than abstracts, the acoustic environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。