低资源翻译中,精选高质量单语数据比堆量更有效。
Quantity vs. Quality of Monolingual Source Data in Automatic Text Translation: Can It Be Too Little If It Is Too Good?
- 按质量或领域相似度筛选单语数据,而非全量使用。
- 在英德低资源翻译任务中,少而精的数据提升模型性能。
- 适合数据稀缺但质量可控的机器翻译场景。
单语数据因数量庞大,常被用于扩充稀缺的平行语料以训练更好的自动翻译模型。自学习方法通过让模型从自身输出中学习来利用此类数据。然而已有研究显示,若可用平行数据极少,过多单语数据反而会损害模型性能。本文研究单语数据是否也可能太少,并探讨基于质量的削减对翻译模型表现的影响。实验表明,在英德低资源神经机器翻译任务中,仅选取最相关、高质量的额外数据,通常优于使用全部可用数据。
原文摘要 · Abstract (English)
Monolingual data, being readily available in large quantities, has been used to upscale the scarcely available parallel data to train better models for automatic translation. Self-learning, where a model is made to learn from its output, is one approach to exploit such data. However, it has been shown that too much of this data can be detrimental to the performance of the model if the available parallel data is comparatively extremely low. In this study, we investigate whether the monolingual data can also be too little and if this reduction, based on quality, has any effect on the performance of the translation model. Experiments have shown that on English-German low-resource NMT, it is often better to select only the most useful additional data, based on quality or closeness to the domain of the test data, than utilizing all of the available data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。