首个多语言医疗语音识别数据集及模型,助力跨语言医患沟通
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
- 构建首个涵盖五种语言的医疗语音识别数据集与端到端模型
- 数据量达全球领先,覆盖多种语境、口音和角色,支持多语言对比研究
- 适合医疗AI、语音识别及跨语言系统开发者使用
医疗领域多语言自动语音识别(ASR)是语音翻译、语音理解及语音助手等下游应用的基础。该技术通过跨越语言障碍提升医患沟通效率,缓解专业人力短缺,尤其在疫情中对诊断与治疗具有重要意义。本文提出MultiMed,首个多语言医疗ASR数据集,以及涵盖小到大规模的端到端医疗ASR模型,支持越南语、英语、德语、法语和中文普通话。据我们所知,MultiMed是当前主流基准中规模最大、覆盖时长最长、录音条件最多、口音最丰富、说话角色最全的医疗语音数据集。同时,本文首次开展医疗ASR多语言研究,包含可复现的基线实验、单语与多语对比分析、注意力编码器-解码器(AED)与混合模型对比,以及语言学分析。还提出了适用于工业界固定参数量限制的实用端到端训练方案。所有代码、数据与模型均已开源:https://github.com/leduckhai/MultiMed/tree/master/MultiMed。
原文摘要 · Abstract (English)
Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and voice-activated assistants. This technology improves patient care by enabling efficient communication across language barriers, alleviating specialized workforce shortages, and facilitating improved diagnosis and treatment, particularly during pandemics. In this work, we introduce MultiMed, the first multilingual medical ASR dataset, along with the first collection of small-to-large end-to-end medical ASR models, spanning five languages: Vietnamese, English, German, French, and Mandarin Chinese. To our best knowledge, MultiMed stands as the world's largest medical ASR dataset across all major benchmarks: total duration, number of recording conditions, number of accents, and number of speaking roles. Furthermore, we present the first multilinguality study for medical ASR, which includes reproducible empirical baselines, a monolinguality-multilinguality analysis, Attention Encoder Decoder (AED) vs Hybrid comparative study and a linguistic analysis. We present practical ASR end-to-end training schemes optimized for a fixed number of trainable parameters that are common in industry settings. All code, data, and models are available online: https://github.com/leduckhai/MultiMed/tree/master/MultiMed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。