提出可迁移的语音对抗攻击方法,提升黑盒场景下模型鲁棒性评估能力。
Transferable Adversarial Attacks against ASR
- 基于时域的可迁移攻击结合可微特征提取器
- 在两个数据集上五种模型均优于基线方法
- 兼顾人耳不可感知性,适合实际部署场景
鉴于自动语音识别(ASR)的广泛应用与研究进展,确保其对微小输入扰动的鲁棒性成为实时应用中的关键问题。以往研究多聚焦于白盒环境下模型全量信息可用的情况,但在实际应用中,完整模型信息通常不可获取。因此,评估黑盒ASR模型的鲁棒性至关重要。本文系统研究了先进ASR模型在真实黑盒攻击下的脆弱性,提出结合两种先进的时域可迁移攻击方法与可微特征提取器,并设计了一种语音感知梯度优化方法(SAGO),通过语音活动检测规则和语音感知梯度优化器,在最小影响人耳感知的前提下强制错误转录。大量实验结果表明,该方法在两个数据库上的五种模型中均显著优于基线方法。
原文摘要 · Abstract (English)
Given the extensive research and real-world applications of automatic speech recognition (ASR), ensuring the robustness of ASR models against minor input perturbations becomes a crucial consideration for maintaining their effectiveness in real-time scenarios. Previous explorations into ASR model robustness have predominantly revolved around evaluating accuracy on white-box settings with full access to ASR models. Nevertheless, full ASR model details are often not available in real-world applications. Therefore, evaluating the robustness of black-box ASR models is essential for a comprehensive understanding of ASR model resilience. In this regard, we thoroughly study the vulnerability of practical black-box attacks in cutting-edge ASR models and propose to employ two advanced time-domain-based transferable attacks alongside our differentiable feature extractor. We also propose a speech-aware gradient optimization approach (SAGO) for ASR, which forces mistranscription with minimal impact on human imperceptibility through voice activity detection rule and a speech-aware gradient-oriented optimizer. Our comprehensive experimental results reveal performance enhancements compared to baseline approaches across five models on two databases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。