arXiv:2409.02712cs.CLcs.LG2024-09中稿 · I2CT 2024被引 5

用跨语言句子表示筛选数据,提升低资源翻译质量

A Data Selection Approach for Enhancing Low Resource Machine Translation Using Cross-Lingual Sentence Representations

  • 用IndicSBERT衡量中英文句义相似度,过滤错误翻译
  • 过滤后英语-马拉地语翻译准确率显著提升
  • 适合低资源语言对的数据清洗与翻译优化

低资源语言对的机器翻译因平行语料和语言资源稀缺而面临挑战。本研究聚焦英语-马拉地语语言对,现有数据噪声严重,影响模型性能。为缓解数据质量问题,提出基于跨语言句子表示的数据过滤方法。利用多语言SBERT模型,通过IndicSBERT评估原文与译文的语义等价性,保留语义一致的翻译,剔除显著偏差的条目。实验表明,经IndicSBERT过滤后,翻译质量显著优于基线。结果证明,跨语言句子表示可有效降低低资源场景下的翻译错误。该方法将多语言句向量模型融入翻译流程,为低资源环境下的机器翻译技术提供新思路,不仅适用于英语-马拉地语,也可推广至其他低资源语言对。

原文摘要 · Abstract (English)

Machine translation in low-resource language pairs faces significant challenges due to the scarcity of parallel corpora and linguistic resources. This study focuses on the case of English-Marathi language pairs, where existing datasets are notably noisy, impeding the performance of machine translation models. To mitigate the impact of data quality issues, we propose a data filtering approach based on cross-lingual sentence representations. Our methodology leverages a multilingual SBERT model to filter out problematic translations in the training data. Specifically, we employ an IndicSBERT similarity model to assess the semantic equivalence between original and translated sentences, allowing us to retain linguistically correct translations while discarding instances with substantial deviations. The results demonstrate a significant improvement in translation quality over the baseline post-filtering with IndicSBERT. This illustrates how cross-lingual sentence representations can reduce errors in machine translation scenarios with limited resources. By integrating multilingual sentence BERT models into the translation pipeline, this research contributes to advancing machine translation techniques in low-resource environments. The proposed method not only addresses the challenges in English-Marathi language pairs but also provides a valuable framework for enhancing translation quality in other low-resource language translation tasks.

低资源翻译数据清洗跨语言表征IndicSBERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。