arXiv:2608.30348cs.SDeess.AS2026-08中稿 · EMNLP

提出新指标衡量语音增强对大模型语音系统的影响,发现音质提升反而可能误导意图识别。

Perceptually Better, Semantically Worse: Measuring Speech Enhancement Impact on LLM-Based Voice Systems

  • 引入输出分歧率(ODR)量化语音增强对大模型意图分类的影响
  • MetricGAN+使意图错误率翻倍(ODR达0.318),远高于未增强噪声的0.135
  • 现有音质评估指标无法预测大模型性能下降,适合关注语音-大模型链路的研究者

语音增强(SE)常作为语音AI流程的预处理步骤,假设更优音质能提升下游任务表现。然而,SE引入的失真是否影响下游大语言模型(LLM)任务性能仍未知。本文提出输出分歧率(ODR),用于衡量在不同增强条件下,语音增强后的输入导致大模型意图分类结果与原始清晰语音相比发生改变的频率。在2,974个SLURP语音片段上,使用Whisper large-v3和wav2vec2-large级联模型进行测试,所有条件的ODR均显著高于零(p < 0.001,二项式检验)。MetricGAN{+}虽提升PESQ得分,但其ODR高达0.318,是未增强噪声的两倍以上(0.135)。未经抑制的回声通过说话人替换导致高达0.836的ODR,这一失败无法被字错率(WER)捕捉。音频质量指标与ODR相关性极低(SQUIM-MOS ρ = -0.068,PESQ ρ = -0.467)。该现象在多种自动语音识别架构中重复验证,表明标准音质指标不足以评估基于大模型的语音系统质量。

原文摘要 · Abstract (English)

Speech enhancement (SE) is commonly applied as a preprocessing step in spoken AI pipelines under the assumption that better audio quality improves downstream task performance. Whether SE-induced distortions propagate to downstream LLM task performance remains an open question. We introduce Output Divergence Rate (ODR), which measures how often SE changes an LLM's intent classification relative to clean speech, and benchmark five conditions on 2,974 SLURP clips using Whisper large-v3 and wav2vec2-large cascades. Every condition produces ODR significantly above zero ($p < 0.001$, binomial test). MetricGAN{+} more than doubles ODR versus unenhanced noisy speech (0.318 vs. 0.135) despite improving PESQ, and unmitigated echo reaches an ODR of 0.836 through speaker substitution, a failure WER cannot capture. Audio quality metrics range from near-zero to moderate correlation with ODR (SQUIM-MOS $ρ=-0.068$, PESQ $ρ=-0.467$). The MetricGAN{+} and echo results replicate across ASR architectures, indicating that standard audio quality metrics are insufficient for LLM pipeline quality.

语音增强大模型评估语音质量意图识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。