用生成式方法对齐阿英法律文学文本,提升翻译质量。
AlignAR: Generative Sentence Alignment for Arabic-English Parallel Corpora of Legal and Literary Texts
- 基于大模型生成句子对齐,突破传统一对一映射限制。
- 在复杂文本上实现85.5%的F1分数,较旧方法提升近9%。
- 适合需要高质量双语数据的研究者与教学应用。
高质量平行语料库对机器翻译研究和翻译教学至关重要。然而,阿拉伯语-英语资源仍十分稀缺,现有数据集主要为简单的一一对应关系。本文提出AlignAR,一种生成式句子对齐方法,并构建了一个包含简单法律文本和复杂文学文本的新阿拉伯语-英语语料库。评估表明,'Easy'数据集缺乏区分能力,无法全面检验对齐方法性能;通过减少'Hard'子集中的一一对应关系,暴露了传统方法的局限性。相比之下,基于大语言模型的方法展现出更强鲁棒性,在整体上达到85.5%的F1分数,较先前方法提升近9%。相关数据集与代码已开源至https://github.com/XXX。
原文摘要 · Abstract (English)
High-quality parallel corpora are essential for Machine Translation (MT) research and translation teaching. However, Arabic-English resources remain scarce and existing datasets mainly consist of simple one-to-one mappings. In this paper, we present AlignAR, a generative sentence alignment method, and a new Arabic-English dataset comprising simple legal and complex literary parallel texts. Our evaluation demonstrates that "Easy" datasets lack the discriminatory power to fully assess alignment methods. By reducing one-to-one mappings in our "Hard" subset, we exposed the limitations of traditional alignment methods. In contrast, LLM-based approaches demonstrated better robustness, achieving an overall F1-score of 85.5%, a nearly 9% improvement over previous methods. Our datasets and codes are open-sourced at https://github.com/XXX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。