用多个语音识别模型的差异,自动定位医疗录音中可能出错的段落。
From Black Box to Glass Box: Cross-Model ASR Disagreement to Prioto Review in Ambient AI Scribe Documentation
- 通过多模型对比,用分歧度识别潜在错误区域。
- 72.1%的字词一致,但2.5%存在高风险分歧,且受口音影响明显。
- 无需人工标注即可发现可疑片段,适合临床文档审核场景。
环境智能“记录员”系统可减轻临床文书负担,但自动语音识别(ASR)错误常因缺乏仔细审查而未被发现,且高质量人工参考转录往往不可得。本文研究是否可利用异构ASR系统间的跨模型分歧,作为无参考的不确定性信号,以优先指引医疗转录中的人工核查。使用50个公开医学教育音频片段(共8小时14分钟),分别用8种不同来源的ASR系统进行转录,包括商业API与开源引擎。对多模型输出进行对齐,构建共识伪参考,并采用多数强度指标量化逐词一致性;进一步按类型(内容/标点/格式)分析分歧,并通过留一法(jackknife)共识评分评估各模型表现。跨模型一致性较低(ICC[2,1] = 0.131),表明各系统故障模式差异显著。在76,398个评估词位中,72.1%显示高度一致(7-8模型一致),2.5%落入高风险区(0-3模型一致),高风险占比在不同口音组间介于0.7%至11.4%。低一致区域中,内容类分歧占比较高,随高风险质量五分位数上升,内容分歧比例从53.9%增至73.9%。结果表明,跨模型分歧可提供稀疏但可定位的异常信号,无需人工参考即可暴露潜在不可靠转录段,支持针对性人工核查;但被标记区域的实际临床准确性仍需验证。
原文摘要 · Abstract (English)
Ambient AI "scribe" systems promise to reduce clinical documentation burden, but automatic speech recognition (ASR) errors can remain unnoticed without careful review, and high-quality human reference transcripts are often unavailable for calibrating uncertainty. We investigate whether cross-model disagreement among heterogeneous ASR systems can act as a reference-free uncertainty signal to prioritize human verification in medical transcription workflows. Using 50 publicly available medical education audio clips (8 h 14 min), we transcribed each clip with eight ASR systems spanning commercial APIs and open-source engines. We aligned multi-model outputs, built consensus pseudo-references, and quantified token-level agreement using a majority-strength metric; we further characterized disagreements by type (content vs. punctuation/formatting) and assessed per-model agreement via leave-one-model-out (jackknife) consensus scoring. Inter-model reliability was low (ICC[2,1] = 0.131), indicating heterogeneous failure modes across systems. Across 76,398 evaluated token positions, 72.1% showed near-unanimous agreement (7-8 models), while 2.5% fell into high-risk bands (0-3 models), with high-risk mass varying from 0.7% to 11.4% across accent groups. Low-agreement regions were enriched for content disagreements, with the content fraction increasing from 53.9% to 73.9% across quintiles of high-risk mass. These results suggest that cross-model disagreement provides a sparse, localizable signal that can surface potentially unreliable transcript spans without human-verified references, enabling targeted review; clinical accuracy of flagged regions remains to be established.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。