arXiv:2511.17153cs.CL2025-11ACL

LangMark为多语言自动译后编辑提供大规模标注数据集。

LangMark: A Multilingual Dataset for Automatic Post-Editing

  • 构建包含20万+三元组的多语言译后编辑数据集。
  • 少样本提示下大模型性能超越商用翻译系统。
  • 适合研究机器翻译纠错与低资源语言优化者。

自动译后编辑(APE)旨在修正机器翻译文本中的错误,提升翻译质量并减少人工干预。尽管神经机器翻译(NMT)取得进展,但有效APE系统的开发仍受限于缺乏大规模、专用于NMT输出的多语言数据集。为此,我们发布并公开了LangMark,一个由专家语言学家人工标注的多语言APE数据集,涵盖英语到巴西葡萄牙语、法语、德语、意大利语、日语、俄语和西班牙语共七种语言。数据集包含206,983个三元组,每个包含源段落、其NMT输出及人工译后编辑版本。利用该数据集,我们实证表明,采用少样本提示的大语言模型(LLMs)能有效执行APE,性能优于领先的商用乃至专有机器翻译系统。我们认为这一新资源将推动未来APE系统的发展与评估。

原文摘要 · Abstract (English)

Automatic post-editing (APE) aims to correct errors in machine-translated text, enhancing translation quality, while reducing the need for human intervention. Despite advances in neural machine translation (NMT), the development of effective APE systems has been hindered by the lack of large-scale multilingual datasets specifically tailored to NMT outputs. To address this gap, we present and release LangMark, a new human-annotated multilingual APE dataset for English translation to seven languages: Brazilian Portuguese, French, German, Italian, Japanese, Russian, and Spanish. The dataset has 206,983 triplets, with each triplet consisting of a source segment, its NMT output, and a human post-edited translation. Annotated by expert human linguists, our dataset offers both linguistic diversity and scale. Leveraging this dataset, we empirically show that Large Language Models (LLMs) with few-shot prompting can effectively perform APE, improving upon leading commercial and even proprietary machine translation systems. We believe that this new resource will facilitate the future development and evaluation of APE systems.

自动译后编辑多语言大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。