arXiv:2508.16048cs.CLcs.AI2025-08中稿 · WMT 2025被引 2

构建首个健康领域多语言平行语料库,助力低资源语言翻译研究

OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages

  • 基于世卫组织官方材料构建2978篇文档级平行语料
  • 大模型在低资源语言上比传统模型提升4.79个ChrF点
  • 文档级上下文对专业领域翻译提升显著,适合医疗翻译研究

机器翻译中,医疗领域因应用广泛且术语专精而具有高风险。然而,该领域低资源语言的评估数据集稀缺。为此,我们推出OpenWHO,一个来自世卫组织在线学习平台的文档级平行语料库,包含2,978篇文档和26,824句,覆盖20多种语言,其中9种为低资源语言。语料源自专家撰写、专业翻译内容,避免网络爬取干扰。利用此资源,我们评估了现代大语言模型(LLMs)与传统机器翻译模型的表现。结果表明,大语言模型持续优于传统模型,其中Gemini 2.5 Flash在低资源测试集上比NLLB-54B高出4.79个ChrF点。进一步分析显示,文档级上下文使用在医疗等专业领域效果最显著。我们已公开发布OpenWHO语料库,以推动低资源医疗翻译研究。

原文摘要 · Abstract (English)

In machine translation (MT), health is a high-stakes domain characterised by widespread deployment and domain-specific vocabulary. However, there is a lack of MT evaluation datasets for low-resource languages in this domain. To address this gap, we introduce OpenWHO, a document-level parallel corpus of 2,978 documents and 26,824 sentences from the World Health Organization's e-learning platform. Sourced from expert-authored, professionally translated materials shielded from web-crawling, OpenWHO spans a diverse range of over 20 languages, of which nine are low-resource. Leveraging this new resource, we evaluate modern large language models (LLMs) against traditional MT models. Our findings reveal that LLMs consistently outperform traditional MT models, with Gemini 2.5 Flash achieving a +4.79 ChrF point improvement over NLLB-54B on our low-resource test set. Further, we investigate how LLM context utilisation affects accuracy, finding that the benefits of document-level translation are most pronounced in specialised domains like health. We release the OpenWHO corpus to encourage further research into low-resource MT in the health domain.

机器翻译医疗翻译低资源语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。