数据中上下文样本少,导致机器翻译难以理解语境。
You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models
- 通过控制数据中上下文样本比例,验证了上下文稀缺是模型瓶颈。
- 改进单一语境现象后,其他语境任务表现不提升,效果不通用。
- 提出两种训练策略,单/多语言场景下翻译准确率分别提升6%和8%。
实现人类水平的翻译需要利用上下文以确保连贯性,并处理代词消歧等复杂现象。标准训练数据中上下文丰富的样本稀疏,被认为是上下文利用困难的原因。本文在单语和多语设置下,系统性地验证了这一假设,通过构建具有可控上下文样本比例的训练数据集。结果表明,训练数据稀疏性与模型性能存在强关联,证实稀疏性是关键瓶颈。值得注意的是,某一语境现象的改进无法泛化到其他现象。尽管存在一定的跨语言迁移,但同语族语言间的迁移优势并不显著。最后,我们提出了两种训练策略并进行实证评估,有效提升了上下文利用率,在ctxPro评测中,单语和多语设置下的准确率分别提高6和8个百分点。
原文摘要 · Abstract (English)
Achieving human-level translations requires leveraging context to ensure coherence and handle complex phenomena like pronoun disambiguation. Sparsity of contextually rich examples in the standard training data has been hypothesized as the reason for the difficulty of context utilization. In this work, we systematically validate this claim in both single- and multilingual settings by constructing training datasets with a controlled proportions of contextually relevant examples. We demonstrate a strong association between training data sparsity and model performance confirming sparsity as a key bottleneck. Importantly, we reveal that improvements in one contextual phenomenon do no generalize to others. While we observe some cross-lingual transfer, it is not significantly higher between languages within the same sub-family. Finally, we propose and empirically evaluate two training strategies designed to leverage the available data. These strategies improve context utilization, resulting in accuracy gains of up to 6 and 8 percentage points on the ctxPro evaluation in single- and multilingual settings respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。