arXiv:2410.03797cs.CL2024-10

用大模型提升印度口音医生语音转录准确率

Searching for Best Practices in Medical Transcription with Large Language Model

  • 结合大模型与语言建模,降低术语识别错误
  • 在医疗录音集上显著降低词错误率,提升关键术语识别
  • 适合需要高精度临床记录的医疗场景

医学独白(尤其是含大量专业术语且带有特定口音)的自动转录对现有系统构成重大挑战。本文提出一种新方法,利用大语言模型(LLM)从医生独白音频中生成高精度医疗转录,特别针对印度口音。该方法融合先进语言建模技术,有效降低词错误率(WER),确保关键医学术语的精准识别。在涵盖多种医学录音的综合数据集上进行严格测试,结果表明该方法在整体转录准确性和关键医学术语保真度方面均有显著提升。研究显示,该系统可显著辅助临床文档工作,为医护人员提供可靠、高精度的转录工具。

原文摘要 · Abstract (English)

The transcription of medical monologues, especially those containing a high density of specialized terminology and delivered with a distinct accent, presents a significant challenge for existing automated systems. This paper introduces a novel approach leveraging a Large Language Model (LLM) to generate highly accurate medical transcripts from audio recordings of doctors' monologues, specifically focusing on Indian accents. Our methodology integrates advanced language modeling techniques to lower the Word Error Rate (WER) and ensure the precise recognition of critical medical terms. Through rigorous testing on a comprehensive dataset of medical recordings, our approach demonstrates substantial improvements in both overall transcription accuracy and the fidelity of key medical terminologies. These results suggest that our proposed system could significantly aid in clinical documentation processes, offering a reliable tool for healthcare providers to streamline their transcription needs while maintaining high standards of accuracy.

语音转录医疗AI大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。