提升语音识别抗干扰能力,让系统更可靠地识别真实语音。
What Was That Again? Certified Robustness for Automatic Speech Recognition

- 用双门诊断流程验证每个词是否存在且不受攻击干扰。
- 在四种模型上实现最高55%的词错误率降低。
- 适合关注语音系统安全与可信度的研究者和开发者。
自动语音识别系统对对抗性及正常扰动均极为敏感。尽管已有参考数据集多次证明此现象,但在部署系统中检测此类行为却极其困难,因缺乏真实转录的权威知识。本文通过引入认证机制,显著降低了词错误率(WER),提升了召回率,并减少了置信度与WER之间的斯皮尔曼相关性。核心方法为双门诊断流水线:双向原子审计通过累积统计证据,认证词的存在性与对抗性排除;基于排名的锦标赛机制选择最优输出序列。在四种不同架构上的评估显示,词错误率相对减少最高达55%,同时提供词级与句级的细粒度认证,增强声学安全性。
原文摘要 · Abstract (English)
Automatic Speech Recognition systems are notoriously both sensitive to adversarial and benign perturbations. While this has been repeatedly demonstrated using reference datasets, detecting such behaviors in deployed systems is incredibly challenging, due to the absence of oracle knowledge of the true transcription. We demonstrate that employing a certification-inspired mechanism can significantly decrease WER, increase recall, and decrease the Spearman correlation between confidence and WER. We achieve this through a dual-gate diagnostic pipeline: a Two-Sided Atomic Audit that accumulates statistical wealth to certify both token existence and adversarial exclusion, and a Rank-Based Tournament that selects the winning sequence. Our evaluations across four diverse architectures demonstrate up to a 55% relative reduction in Word Error Rate, while also providing granular word- and sentence-level certifications to enhance acoustic security.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。