ASR错误影响临床理解,现有评估方法无法准确捕捉其风险。
WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue
- 用专家标注临床对话差异的严重性,建立真实评估基准
- 发现传统指标如WER与临床风险关联度极低
- 用大模型优化评判系统,实现接近专家水平的自动化评估
随着自动语音识别(ASR)在临床对话中日益应用,评估仍主要依赖词错误率(WER)。本文挑战这一标准,探究WER及其他常见指标是否与转录错误的临床影响相关。通过让临床专家对比真实语句与ASR生成结果,对两个医生-患者对话数据集中的差异进行临床影响标注(无、轻微或显著)。分析显示,WER及一系列现有指标与专家评定的风险标签相关性很差。为弥补评估空白,本文提出基于大模型的评判框架,使用GEPA和DSPy对Gemini-2.5-Pro进行程序化优化,使其达到人类水平:准确率达90%,科恩κ系数达0.816。该研究提供了一个可验证、自动化的评估体系,推动ASR评估从文本保真度转向临床安全性的可扩展评估。
原文摘要 · Abstract (English)
As Automatic Speech Recognition (ASR) is increasingly deployed in clinical dialogue, standard evaluations still rely heavily on Word Error Rate (WER). This paper challenges that standard, investigating whether WER or other common metrics correlate with the clinical impact of transcription errors. We establish a gold-standard benchmark by having expert clinicians compare ground-truth utterances to their ASR-generated counterparts, labeling the clinical impact of any discrepancies found in two distinct doctor-patient dialogue datasets. Our analysis reveals that WER and a comprehensive suite of existing metrics correlate poorly with the clinician-assigned risk labels (No, Minimal, or Significant Impact). To bridge this evaluation gap, we introduce an LLM-as-a-Judge, programmatically optimized using GEPA through DSPy to replicate expert clinical assessment. The optimized judge (Gemini-2.5-Pro) achieves human-comparable performance, obtaining 90% accuracy and a strong Cohen's kappa of 0.816. This work provides a validated, automated framework for moving ASR evaluation beyond simple textual fidelity to a necessary, scalable assessment of safety in clinical dialogue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。