arXiv:2505.22251eess.AScs.CL2025-05被引 10

LLM语音识别评估常被数据污染误导,真实性能被高估。

Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition

  • 用含测试集的训练数据训练LLM,导致模型‘见过’测试内容
  • 污染的LLM对测试句生成概率显著更高,但错误率差异微小
  • 提醒研究者必须用未参与训练的数据评估模型

近期研究声称大语言模型(LLMs)可提升语音任务性能,常引用LibriSpeech和Common Voice数据集结果。但本研究发现,这两个数据集的大量内容出现在公开的LLM预训练语料中。为评估污染影响,对比了含/不含污染训练的LLM。含污染的LLM更可能生成训练中见过的测试句。基于这些LLM的语音识别系统显示,错误率差异微小,但对训练中出现过的转录文本赋予显著更高的概率。结果表明,哪怕少量数据污染也会导致LLM输出偏差,强调必须使用保留数据评估基于LLM的语音系统。

原文摘要 · Abstract (English)

Recent work suggests that large language models (LLMs) can improve performance of speech tasks compared to existing systems. To support their claims, results on LibriSpeech and Common Voice are often quoted. However, this work finds that a substantial amount of the LibriSpeech and Common Voice evaluation sets appear in public LLM pretraining corpora. This calls into question the reliability of findings drawn from these two datasets. To measure contamination impact, LLMs trained with/without contamination are compared. A contaminated LLM is more likely to generate test sentences it has seen during training. Then, speech recognisers based on LLMs are compared. They show only subtle error rate differences if the LLM is contaminated, but assign significantly higher probabilities to transcriptions seen during LLM training. Results show that LLM outputs can be biased by tiny amounts of data contamination, highlighting the importance of evaluating LLM-based speech systems with held-out data.

语音识别LLM数据污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。