通过提取语音识别中的口音特征子空间,揭示其与模型性能的深层关联。
ACES: Accent Subspaces for Coupling, Explanations, and Stress-Testing in Automatic Speech Recognition
- 从语音识别表征中提取口音判别子空间
- 口音子空间扰动使错误率差距扩大近50%(21.3→31.8个百分点)
- 适合关注语音识别公平性与鲁棒性的研究者
语音识别系统在不同口音间存在持续的性能差异,但这些差距是表面偏见还是深层结构脆弱性尚不明确。我们提出ACES,一种三阶段审计方法:从语音识别表征中提取口音判别子空间,将对抗攻击限制于该子空间,并测试移除它是否提升公平性。在Wav2Vec2-base模型上针对七种口音进行实验,沿口音子空间施加近乎不可察觉的扰动(约60 dB SNR),使错误率差距扩大近50%(21.3→31.8个百分点),超过随机子空间对照;置换标签测试确认其对真实口音结构的特异性。部分移除该子空间反而恶化了错误率与差距,表明口音判别特征与识别关键特征深度纠缠。因此,口音子空间可作为强大的公平性审计工具,而非简单的消除手段。
原文摘要 · Abstract (English)
ASR systems exhibit persistent performance disparities across accents, but whether these gaps reflect superficial biases or deep structural vulnerabilities remains unclear. We introduce ACES, a three-stage audit that extracts accent-discriminative subspaces from ASR representations, constrains adversarial attacks to them, and tests whether removing them improves fairness. On Wav2Vec2-base with seven accents, imperceptible perturbations (~60 dB SNR) along the accent subspace amplify the WER disparity gap by nearly 50% (21.3->31.8 pp), exceeding random-subspace controls; a permuted-label test confirms specificity to genuine accent structure. Partially removing the subspace worsens both WER and disparity, revealing that accent-discriminative and recognition-critical features are deeply entangled. ACES thus positions accent subspaces as powerful fairness-auditing tools, not simple erasure levers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。