arXiv:2502.19074cs.CL2025-02EMNLP被引 2

用去偏启发式方法提升低资源语言平行语料质量

Improving the quality of Web-mined Parallel Corpora of Low-Resource Languages using Debiasing Heuristics

  • 通过启发式规则过滤多语言模型产生的偏差噪声
  • 训练出的NMT模型性能提升且跨模型差异减小
  • 适合低资源语言数据清洗与NMT研究者使用

平行数据筛选(PDC)旨在从网络挖掘的平行语料中去除噪声句子。目前主流方法是利用预训练多语言模型(multiPLM)的句向量相似度对句子对进行排序。然而,已有研究表明,不同multiPLM的选择显著影响筛选后语料质量,由此训练的神经机器翻译(NMT)模型在性能上存在差异。本文揭示该差异源于不同multiPLM对特定类型句子对存在偏差,这些句子在NMT视角下被误判为噪声。我们提出一系列启发式方法,可有效移除此类噪声。基于清理后的语料训练的NMT模型表现更优,且跨multiPLM的性能差异显著降低。论文公开了源代码和清理后的数据集。

原文摘要 · Abstract (English)

Parallel Data Curation (PDC) techniques aim to filter out noisy parallel sentences from web-mined corpora. Ranking sentence pairs using similarity scores on sentence embeddings derived from Pre-trained Multilingual Language Models (multiPLMs) is the most common PDC technique. However, previous research has shown that the choice of the multiPLM significantly impacts the quality of the filtered parallel corpus, and the Neural Machine Translation (NMT) models trained using such data show a disparity across multiPLMs. This paper shows that this disparity is due to different multiPLMs being biased towards certain types of sentence pairs, which are treated as noise from an NMT point of view. We show that such noisy parallel sentences can be removed to a certain extent by employing a series of heuristics. The NMT models, trained using the curated corpus, lead to producing better results while minimizing the disparity across multiPLMs. We publicly release the source code and the curated datasets.

数据清洗低资源语言NMT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。