构建学术翻译平行语料库,提升论文跨语言理解与生成质量
ACADATA: Parallel Dataset of Academic Data for Machine Translation
- 构建150万组作者自撰学术段落对,覆盖96种语言方向
- 微调后7B模型在学术翻译上提升12.4 d-BLEU,长文本翻译效率提高24.9%
- 适合学术翻译、跨语言研究者及大模型优化团队使用
我们提出ACADATA,一个高质量的学术翻译平行数据集,包含两个子集:ACAD-TRAIN包含约150万组作者生成的段落对,覆盖96种语言方向;ACAD-BENCH为精选的近6,000条翻译评估集,涵盖12种语言方向。为验证其有效性,我们在ACAD-TRAIN上微调两个大语言模型(LLM),并在ACAD-BENCH上与专用机器翻译系统、通用开源模型及多个大规模专有模型进行对比。实验结果表明,在ACAD-TRAIN上微调可使7B和2B模型在学术翻译上的平均性能分别提升6.1和12.4 d-BLEU点,同时在英文输出长文本翻译任务中整体效率提升最高达24.9%。最佳微调模型超越了当前最优的专有与开源模型。通过发布ACAD-TRAIN、ACAD-BENCH及微调模型,我们为学术领域和长文本翻译研究提供重要资源。
原文摘要 · Abstract (English)
We present ACADATA, a high-quality parallel dataset for academic translation, that consists of two subsets: ACAD-TRAIN, which contains approximately 1.5 million author-generated paragraph pairs across 96 language directions and ACAD-BENCH, a curated evaluation set of almost 6,000 translations covering 12 directions. To validate its utility, we fine-tune two Large Language Models (LLMs) on ACAD-TRAIN and benchmark them on ACAD-BENCH against specialized machine-translation systems, general-purpose, open-weight LLMs, and several large-scale proprietary models. Experimental results demonstrate that fine-tuning on ACAD-TRAIN leads to improvements in academic translation quality by +6.1 and +12.4 d-BLEU points on average for 7B and 2B models respectively, while also improving long-context translation in a general domain by up to 24.9% when translating out of English. The fine-tuned top-performing model surpasses the best propietary and open-weight models on academic translation domain. By releasing ACAD-TRAIN, ACAD-BENCH and the fine-tuned models, we provide the community with a valuable resource to advance research in academic domain and long-context translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。