将大模型与音频处理结合,提升医疗语音识别准确率。
Au-M-ol: A Unified Model for Medical Audio and Language Understanding

- 用音频编码器提取医疗语音特征,映射到大模型输入空间。
- 在医疗语音转写任务上,错误率降低56%。
- 适合临床场景中复杂环境下的语音理解需求。
本文提出Au-M-ol,一种融合大语言模型(LLM)与音频处理的多模态架构,旨在提升医疗语音识别等临床任务性能。该模型包含三个核心组件:(1) 音频编码器,用于从医疗语音中提取丰富声学特征;(2) 适配层,将音频特征映射至大模型输入空间;(3) 预训练大语言模型,执行转录与临床语言理解。此设计使模型可直接解析口语化医疗内容,显著提升准确性与鲁棒性。实验表明,相较于现有最佳基线,Au-M-ol在医疗语音转写任务上的词错误率(WER)降低56%。模型在噪声环境、专业术语及说话人差异等挑战条件下仍表现优异,表明其在真实临床应用中具备潜力,尤其适用于对可靠、上下文感知语音理解有高要求的场景。
原文摘要 · Abstract (English)
In this work, we present Au-M-ol, a novel multimodal architecture that extends Large Language Models (LLMs) with audio processing. It is designed to improve performance on clinically relevant tasks such as Automatic Speech Recognition (ASR). Au-M-ol has three main components: (1) an audio encoder that extracts rich acoustic features from medical speech, (2) an adaptation layer that maps audio features into the LLM input space, and (3) a pretrained LLM that performs transcription and clinical language understanding. This design allows the model to interpret spoken medical content directly, improving both accuracy and robustness. In experiments, Au-M-ol reduces Word Error Rate (WER) by 56\% compared to state-of-the-art baselines on medical transcription tasks. The model also performs well in challenging conditions, including noisy environments, domain-specific terminology, and speaker variability. These results suggest that Au-M-ol is a strong candidate for real-world clinical applications, where reliable and context-aware audio understanding is essential.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。