arXiv:2602.12249cs.AIcs.CL2026-02

语音模型在真实场景中常误转街道名,尤其对非英语母语者影响更大。

"Sorry, I Didn't Catch That": How Speech Models Miss What Matters Most

  • 用合成语音生成多样化发音数据,提升模型对非英语母语者的识别能力
  • 非英语母语者误转导致的导航距离误差是英语母语者的两倍
  • 仅用不到1000个合成样本,就能让非英语母语者识别准确率提升近60%

尽管语音识别系统在标准基准上达到低词错误率,但在真实部署中对短而高风险的语句仍频繁失败。本文研究了美国街道名在真实场景下的转录问题,评估了来自OpenAI、Deepgram、Google和Microsoft的15个模型,使用来自语言多样化的美国说话人录音,发现平均转录错误率达44%。我们量化了错误转录对地理定位的影响,发现所有说话人均受影响,但非英语母语者因误转导致的导航距离误差是英语母语者的两倍。为缓解此问题,我们提出一种基于开源语音合成模型的合成数据生成方法,通过不足1000个合成样本微调,使非英语母语者在街道名转录上的准确率相对基线提升近60%。结果揭示了基准性能与实际可靠性之间的关键差距,并展示了减少高风险转录错误的简单可扩展路径。

原文摘要 · Abstract (English)

Despite speech recognition systems achieving low word error rates on standard benchmarks, they often fail on short, high-stakes utterances in real-world deployments. Here, we study this failure mode in a high-stakes task: the transcription of U.S. street names as spoken by U.S. participants. We evaluate 15 models from OpenAI, Deepgram, Google, and Microsoft on recordings from linguistically diverse U.S. speakers and find an average transcription error rate of 44%. We quantify the downstream impact of failed transcriptions by geographic locations and show that mis-transcriptions systematically cause errors for all speakers, but that routing distance errors are twice as large for non-English primary speakers compared to English primary speakers. To mitigate this harm, we introduce a synthetic data generation approach that produces diverse pronunciations of named entities using open-source text-to-speech models. Fine-tuning with less than 1,000 synthetic samples improves street name transcription accuracy by nearly 60% (relative to base models) for non-English primary speakers. Our results highlight a critical gap between benchmark performance and real-world reliability in speech systems and demonstrate a simple, scalable path to reducing high-stakes transcription errors.

语音识别公平性合成数据高风险应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。