用半自动方法构建心脏病学领域中英术语库,提升翻译质量与可控性。
Creating Domain-Specific Translation Memories for Machine Translation Fine-tuning: The TRENCARD Bilingual Cardiology Corpus
- 利用译者常用工具,半自动构建领域术语库,确保数据质量与控制权。
- 建成含5万句、约80万词的心脏病学中英双语语料库(TRENCARD)。
- 适合医疗翻译、机器翻译微调及大模型训练,助力专业领域精准翻译。
本文研究如何由译者或其他语言专业人士创建术语库(TM),以构建特定领域的平行语料,用于机器翻译训练与微调、术语库复用以及大语言模型微调等场景。提出一种半自动术语库构建方法,主要依赖译者日常使用的翻译工具,以保障数据质量并实现译者对数据的自主控制。该方法被应用于从土耳其心脏病学期刊的双语摘要中构建以心脏病学为核心的土-英语料库。最终形成的TRENCARD语料库包含约80万源词和5万句子。通过此方法,译者可在合理时间内构建个性化术语库,并用于各类双语任务。
原文摘要 · Abstract (English)
This article investigates how translation memories (TM) can be created by translators or other language professionals in order to compile domain-specific parallel corpora , which can then be used in different scenarios, such as machine translation training and fine-tuning, TM leveraging, and/or large language model fine-tuning. The article introduces a semi-automatic TM preparation methodology leveraging primarily translation tools used by translators in favor of data quality and control by the translators. This semi-automatic methodology is then used to build a cardiology-based Turkish -> English corpus from bilingual abstracts of Turkish cardiology journals. The resulting corpus called TRENCARD Corpus has approximately 800,000 source words and 50,000 sentences. Using this methodology, translators can build their custom TMs in a reasonable time and use them in their bilingual data requiring tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。