arXiv:2501.16533cs.CLcs.LG2025-01

对比三种数据过滤方法在生物医学英波翻译中的效果,推荐LASER提升翻译质量。

A comparison of data filtering techniques for English-Polish LLM-based machine translation in the biomedical domain

  • 用LASER、MUSE、LaBSE过滤医学语料,缩小数据集提升效率。
  • 使用mBART50模型在不同规模数据上微调,测试集上获得最高SacreBLEU得分。
  • LASER表现最优,生成翻译更流畅自然,适合医学领域应用。

大语言模型(LLMs)已成为机器翻译的最先进方法,通常在从网络抓取的海量双语平行语料上训练,这些语料包含低质量条目和冗余信息,带来显著计算挑战。已有多种数据过滤方法用于减少数据集规模,但其有效性因语言对和领域而异。本文评估了常见的数据过滤技术(如LASER、MUSE、LaBSE)在英语-波兰语生物医学领域的翻译表现。通过过滤UFAL医学语料库,构建不同规模的数据集以微调mBART50模型,并在Khresmoi数据集上使用SacreBLEU指标进行评估,同时由双语母语者评估译文质量。结果表明,LASER与MUSE均可显著缩减数据量,同时保持甚至提升翻译性能。我们推荐使用LASER,因其持续优于其他方法,且生成译文最为流畅自然。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become state-of-the-art in Machine Translation (MT), often trained on massive bilingual parallel corpora scraped from the web, that contain low-quality entries and redundant information, leading to significant computational challenges. Various data filtering methods exist to reduce dataset sizes, but their effectiveness largely varies based on specific language pairs and domains. This paper evaluates the impact of commonly used data filtering techniques, such as LASER, MUSE, and LaBSE, on English-Polish translation within the biomedical domain. By filtering the UFAL Medical Corpus, we created varying dataset sizes to fine-tune the mBART50 model, which was then evaluated using the SacreBLEU metric on the Khresmoi dataset, having the quality of translations assessed by bilingual speakers. Our results show that both LASER and MUSE can significantly reduce dataset sizes while maintaining or even enhancing performance. We recommend the use of LASER, as it consistently outperforms the other methods and provides the most fluent and natural-sounding translations.

机器翻译数据过滤生物医学大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。