arXiv:2411.19096cs.CL2024-11被引 1

为印度语种设计新对齐方法,提升文档级机器翻译效果

Pralekha: Cross-Lingual Document Alignment for Indic Languages

  • 用分块匹配替代整体池化,提升对齐精度
  • 新指标DAC使对齐速度提升2-3倍,性能更优
  • 公开超300万对齐文档数据集,支持后续研究

文档级机器翻译的并行文档对挖掘仍面临挑战,现有跨语言文档对齐技术受限于元数据稀少或全局表示无法捕捉细粒度对齐线索。句嵌入模型上下文窗口有限,而基于句子的对齐方式计算开销巨大。为此,我们提出Pralekha基准,包含11种印度语与英语之间的超过300万对齐文档对(其中英印对达150万)。同时引入文档对齐系数(DAC),通过匹配小块文本并以对齐块数占平均块数的比例衡量相似性,实现细粒度对齐。内在评估显示,该方法速度比池化法快2-3倍,性能相当;外在评估表明,基于DAC对齐数据训练的文档级翻译模型持续优于基线方法。结果验证了DAC在并行文档挖掘中的有效性。数据集与评估框架已公开,供后续研究使用。

原文摘要 · Abstract (English)

Mining parallel document pairs for document-level machine translation (MT) remains challenging due to the limitations of existing Cross-Lingual Document Alignment (CLDA) techniques. Existing methods often rely on metadata such as URLs, which are scarce, or on pooled document representations that fail to capture fine-grained alignment cues. Moreover, the limited context window of sentence embedding models hinders their ability to represent document-level context, while sentence-based alignment introduces a combinatorially large search space, leading to high computational cost. To address these challenges for Indic languages, we introduce Pralekha, a benchmark containing over 3 million aligned document pairs across 11 Indic languages and English, which includes 1.5 million English-Indic pairs. Furthermore, we propose Document Alignment Coefficient (DAC), a novel metric for fine-grained document alignment. Unlike pooling-based methods, DAC aligns documents by matching smaller chunks and computes similarity as the ratio of aligned chunks to the average number of chunks in a pair. Intrinsic evaluation shows that our chunk-based method is 2-3x faster while maintaining competitive performance, and that DAC achieves substantial gains over pooling-based baselines. Extrinsic evaluation further demonstrates that document-level MT models trained on DAC-aligned pairs consistently outperform those using baseline alignment methods. These results highlight DAC's effectiveness for parallel document mining. The dataset and evaluation framework are publicly available to support further research.

文档对齐印度语言机器翻译多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。