arXiv:2601.19945cs.CLcs.AI2026-01

评测29个语音识别模型在德语医疗对话中的表现,发现顶尖模型错误率低于3%。

Benchmarking von ASR-Modellen im deutschen medizinischen Kontext: Eine Leistungsanalyse anhand von Anamnesegesprächen

  • 构建模拟医患对话数据集,测试29个开源与商用语音识别模型。
  • 最佳模型在医疗术语上错误率低于3%,部分模型对方言识别差。
  • 适合医疗AI研发者、语音识别优化者参考,尤其关注德语方言场景。

自动语音识别(ASR)有望显著减轻医护人员负担,例如通过自动化病历记录。尽管英语领域已有诸多基准测试,但针对德语医疗场景的评估仍不足,尤其缺乏对方言影响的考量。本文构建了一个模拟医患对话的精选数据集,评估了共计29个ASR模型,涵盖Whisper、Voxtral、Wav2Vec2等开源模型及AssemblyAI、Deepgram等商业API。采用WER、CER、BLEU三种指标进行评估,并开展定性语义分析展望。结果表明模型性能差异显著:最优系统在医疗术语上已实现低于3%的词错误率(WER),而部分模型在医疗术语或受方言影响的表达上错误率明显更高。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) offers significant potential to reduce the workload of medical personnel, for example, through the automation of documentation tasks. While numerous benchmarks exist for the English language, specific evaluations for the German-speaking medical context are still lacking, particularly regarding the inclusion of dialects. In this article, we present a curated dataset of simulated doctor-patient conversations and evaluate a total of 29 different ASR models. The test field encompasses both open-weights models from the Whisper, Voxtral, and Wav2Vec2 families as well as commercial state-of-the-art APIs (AssemblyAI, Deepgram). For evaluation, we utilize three different metrics (WER, CER, BLEU) and provide an outlook on qualitative semantic analysis. The results demonstrate significant performance differences between the models: while the best systems already achieve very good Word Error Rates (WER) of partly below 3%, the error rates of other models, especially concerning medical terminology or dialect-influenced variations, are considerably higher.

语音识别医疗AI德语Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。