构建科学领域多语种语料库,提升科研文献翻译质量
Enhancing Scientific Discourse: Machine Translation for the Scientific Domain
- 针对癌症、能源等四大领域构建多语言平行语料库
- 在西班牙语-英语等三组语言上实现翻译性能提升
- 适合跨语言科研交流与机器翻译研究者使用
科学研究的快速增长要求跨越语言障碍进行有效沟通。机器翻译(MT)为获取国际出版物提供了可行方案,但科学领域因专业术语密集和句式复杂带来独特挑战。本文构建了面向西班牙语-英语、法语-英语、葡萄牙语-英语三组语言对的科学领域平行与单语语料库。每个语料库包含通用科学语料及聚焦癌症研究、能源研究、神经科学、交通研究四个细分领域的子语料库。通过微调通用神经机器翻译系统,评估语料库质量。文中详细说明了语料构建流程、微调策略,并给出评估结果。
原文摘要 · Abstract (English)
The increasing volume of scientific research necessitates effective communication across language barriers. Machine translation (MT) offers a promising solution for accessing international publications. However, the scientific domain presents unique challenges due to its specialized vocabulary and complex sentence structures. In this paper, we present the development of a collection of parallel and monolingual corpora for the scientific domain. The corpora target the language pairs Spanish-English, French-English, and Portuguese-English. For each language pair, we create a large general scientific corpus as well as four smaller corpora focused on the domains of: Cancer Research, Energy Research, Neuroscience, and Transportation research. To evaluate the quality of these corpora, we utilize them for fine-tuning general-purpose neural machine translation (NMT) systems. We provide details regarding the corpus creation process, the fine-tuning strategies employed, and we conclude with the evaluation results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。