医学翻译中保留英文术语,提升专业准确性。
On Creating an English-Thai Code-switched Machine Translation in Medical Domain
- 构建英泰混用医学翻译数据集,训练混合语言模型。
- 在自动评分与人工评估中均优于GPT-3.5、GPT-4等基线。
- 医生更倾向保留关键英文术语的翻译,即使略显不流畅。
医疗领域机器翻译对提升医疗质量与传播医学知识至关重要。尽管英泰机器翻译技术已有进展,但通用方法常因无法精准翻译医学术语而表现不佳。本研究不仅关注提升翻译准确率,更强调在译文中保留关键英文医学术语,采用代码切换(Code-Switched, CS)翻译策略。我们构建了用于医疗领域的英泰混用翻译数据集,基于该数据微调模型,并与Google NMT及GPT-3.5/GPT-4等强基线进行对比评估。结果表明,该模型在自动指标上表现具有竞争力,在人类偏好评估中得分更高。进一步分析显示,医疗从业者显著偏好保留重要英文术语的翻译,即便其流畅性略有下降。相关代码与测试集已公开于https://github.com/preceptorai-org/NLLB_CS_EM_NLP2024。
原文摘要 · Abstract (English)
Machine translation (MT) in the medical domain plays a pivotal role in enhancing healthcare quality and disseminating medical knowledge. Despite advancements in English-Thai MT technology, common MT approaches often underperform in the medical field due to their inability to precisely translate medical terminologies. Our research prioritizes not merely improving translation accuracy but also maintaining medical terminology in English within the translated text through code-switched (CS) translation. We developed a method to produce CS medical translation data, fine-tuned a CS translation model with this data, and evaluated its performance against strong baselines, such as Google Neural Machine Translation (NMT) and GPT-3.5/GPT-4. Our model demonstrated competitive performance in automatic metrics and was highly favored in human preference evaluations. Our evaluation result also shows that medical professionals significantly prefer CS translations that maintain critical English terms accurately, even if it slightly compromises fluency. Our code and test set are publicly available https://github.com/preceptorai-org/NLLB_CS_EM_NLP2024.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。