通过语音隐空间插值生成自然语音攻击,高效测试语音识别系统漏洞。
Generative Testing of Automated Speech Recognition Systems

- 在文本转语音模型的音素级隐空间中插值,生成自然且具破坏性的语音输入。
- 黑箱攻击成功率98%,失真更低,人评感知质量更高。
- 无需梯度信息,适合评估真实场景下语音系统的鲁棒性。
自动语音识别(ASR)系统已借助基于Transformer的模型实现高精度,广泛应用于关键场景。然而,它们仍易受对抗性攻击影响,尤其在黑箱设置下,攻击需保持听觉自然性。本文提出GATAS,一种基于文本到语音模型音素级隐空间的黑箱测试方法。该方法通过隐表示插值诱导识别错误,同时保持在自然语音流形内。攻击被建模为多目标优化问题,权衡语义偏离与听觉质量。实证评估显示,相较于白盒与黑盒基线,GATAS实现98%的成功率,失真更低,人评感知质量更优。尽管无梯度访问,其性能仍媲美白盒方法,表明表征对齐与感知一致性比内部模型访问更为关键。结果表明,无目标隐空间优化可高效生成真实且有效的ASR测试用例。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) systems have achieved high accuracy with transformer-based models, enabling deployment in critical applications. However, they remain vulnerable to adversarial manipulation, particularly in black-box settings where attacks must preserve perceptual naturalness. This work introduces GATAS, a black-box testing approach that generates failure inducing inputs by operating in the phoneme-level latent space of a text- to-speech model. Instead of perturbing waveforms directly, the approach interpolates latent representations to induce transcription errors while remaining within the manifold of natural speech. The attack is formulated as a multi-objective optimization problem balancing semantic divergence and perceptual quality. Our empirical evaluation against both white-box and black-box baselines shows that GATAS achieves a 98% success rate while producing lower distortion and higher perceptual quality, as confirmed by human studies. Despite operating without gradient access, GATAS remains competitive against white-box methods, highlighting that representation and perceptual alignment are more critical than access to model internals. Overall, our results demonstrate that untargeted latent-space optimization enables the efficient generation of realistic and effective test cases for ASR systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。