针对内镜场景优化语音识别,显著提升准确率与临床可用性。
Development and multi-center evaluation of domain-adapted speech recognition for human-AI teaming in real-world gastrointestinal endoscopy
- 用合成报告分两阶段适配领域术语和噪声环境。
- 字符错误率降为14.14%,医学术语准确率达87.59%。
- 实时性极强,适合边缘设备部署,医生-AI协作更高效。
自动语音识别(ASR)是胃肠镜检查中人机协作的关键接口,但受限于专业术语和复杂声学环境,实际临床应用可靠性不足。本文提出面向真实场景的端到端内镜语音识别系统EndoASR,采用基于合成内镜报告的两阶段适配策略,分别优化领域语言建模与噪声鲁棒性。在六名内镜医师的回顾性评估中,字符错误率(CER)从20.52%降至14.14%,医学术语准确率(Med ACC)从54.30%提升至87.59%。在覆盖五个独立中心的前瞻性多中心研究中,相较于基线模型Paraformer,CER由16.20%降至14.97%,Med ACC由61.63%升至84.16%,验证了其在异构真实场景下的泛化能力。系统实时因子(RTF)仅为0.005,远优于Whisper-large-v3的0.055,模型参数量仅220M,支持高效边缘部署。结合大语言模型后,高质量语音识别显著提升了结构化信息抽取与医患-智能体交互效果。结果表明,领域自适应的语音识别可作为胃肠镜检查中可靠的人机协同接口。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) is a critical interface for human-AI interaction in gastrointestinal endoscopy, yet its reliability in real-world clinical settings is limited by domain-specific terminology and complex acoustic conditions. Here, we present EndoASR, a domain-adapted ASR system designed for real-time deployment in endoscopic workflows. We develop a two-stage adaptation strategy based on synthetic endoscopy reports, targeting domain-specific language modeling and noise robustness. In retrospective evaluation across six endoscopists, EndoASR substantially improves both transcription accuracy and clinical usability, reducing character error rate (CER) from 20.52% to 14.14% and increasing medical term accuracy (Med ACC) from 54.30% to 87.59%. In a prospective multi-center study spanning five independent endoscopy centers, EndoASR demonstrates consistent generalization under heterogeneous real-world conditions. Compared with the baseline Paraformer model, CER is reduced from 16.20% to 14.97%, while Med ACC is improved from 61.63% to 84.16%, confirming its robustness in practical deployment scenarios. Notably, EndoASR achieves a real-time factor (RTF) of 0.005, significantly faster than Whisper-large-v3 (RTF 0.055), while maintaining a compact model size of 220M parameters, enabling efficient edge deployment. Furthermore, integration with large language models demonstrates that improved ASR quality directly enhances downstream structured information extraction and clinician-AI interaction. These results demonstrate that domain-adapted ASR can serve as a reliable interface for human-AI teaming in gastrointestinal endoscopy, with consistent performance validated across multi-center real-world clinical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。