用翻译模型内部特征区分中英文翻译真伪,提升过滤准确率。
Automatic Machine Translation Detection Using a Surrogate Multilingual Translation Model
- 用替代多语言翻译模型的内部表示识别机器翻译
- 非英语语对准确率提升至少5个百分点
- 适合需要高质量训练数据的机器翻译研究者
现代机器翻译系统依赖大量平行语料,这些语料常来自网络,但近期研究表明,其中相当一部分是机器生成的翻译,过度依赖此类合成内容会显著降低翻译质量。因此,过滤非人工翻译已成为构建高质量翻译系统的重要预处理步骤。本文提出一种新方法,直接利用替代多语言翻译模型的内部表示来区分人类与机器翻译句子。实验表明,该方法优于现有最先进技术,尤其在非英语语对上表现突出,准确率提升至少5个百分点。
原文摘要 · Abstract (English)
Modern machine translation (MT) systems depend on large parallel corpora, often collected from the Internet. However, recent evidence indicates that (i) a substantial portion of these texts are machine-generated translations, and (ii) an overreliance on such synthetic content in training data can significantly degrade translation quality. As a result, filtering out non-human translations is becoming an essential pre-processing step in building high-quality MT systems. In this work, we propose a novel approach that directly exploits the internal representations of a surrogate multilingual MT model to distinguish between human and machine-translated sentences. Experimental results show that our method outperforms current state-of-the-art techniques, particularly for non-English language pairs, achieving gains of at least 5 percentage points of accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。