构建首个大规模多语言医疗语音翻译数据集,支持五语种双向互译。
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
- 构建涵盖五语种的多对多医疗语音翻译数据集,覆盖29万条样本。
- 首次系统分析多语言医疗语音翻译,验证端到端优于级联模型。
- 数据与模型开源,适合医疗AI、多语言语音处理研究者使用。
多语言语音翻译(ST)和机器翻译(MT)在医疗领域的应用可提升跨语言沟通效率,缓解专业人才短缺,并在疫情等场景中促进精准诊断与治疗。本文首次系统性研究医疗语音翻译,发布大型医疗语音翻译数据集MultiMed-ST,覆盖越南语、英语、德语、法语及简体/繁体中文五种语言,支持所有翻译方向,共含29万个样本,是目前最大规模的医疗机器翻译数据集和跨领域最大的多对多多语言语音翻译数据集。其次,我们进行了迄今为止最全面的语音翻译分析,包括基准实验、双语与多语言对比、端到端与级联模型对比、任务专用与多任务序列到序列模型对比、语码转换分析以及定量定性错误分析。所有代码、数据与模型均已公开:https://github.com/leduckhai/MultiMed-ST。
原文摘要 · Abstract (English)
Multilingual speech translation (ST) and machine translation (MT) in the medical domain enhances patient care by enabling efficient communication across language barriers, alleviating specialized workforce shortages, and facilitating improved diagnosis and treatment, particularly during pandemics. In this work, we present the first systematic study on medical ST, to our best knowledge, by releasing MultiMed-ST, a large-scale ST dataset for the medical domain, spanning all translation directions in five languages: Vietnamese, English, German, French, and Simplified/Traditional Chinese, together with the models. With 290,000 samples, this is the largest medical MT dataset and the largest many-to-many multilingual ST among all domains. Secondly, we present the most comprehensive ST analysis in the field's history, to our best knowledge, including: empirical baselines, bilingual-multilingual comparative study, end-to-end vs. cascaded comparative study, task-specific vs. multi-task sequence-to-sequence comparative study, code-switch analysis, and quantitative-qualitative error analysis. All code, data, and models are available online: https://github.com/leduckhai/MultiMed-ST
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。