用质量评估反馈提升翻译中代词的准确率。
Quality Estimation based Feedback Training for Improving Pronoun Translation
- 通过质量评估模型生成反馈信号,迭代优化预训练翻译模型。
- 在多个数据集上显著提升代词翻译准确率和整体翻译质量。
- 无需人工标注,适合需要上下文理解的翻译场景。
代词翻译是神经机器翻译中的长期挑战,通常需要句间上下文以保证语言准确性。为此,我们提出 ProNMT 框架,旨在增强上下文感知机器翻译系统中代词及整体翻译质量。ProNMT 利用质量评估(QE)模型与基于代词生成概率的反馈机制,无需依赖大量人工标注,即可迭代微调预训练的 NMT 模型。该框架结合 QE 分数与代词特异性奖励指导训练,确保对语言细微差别的更好处理。大量实验表明,在多个指标下,代词翻译准确率和整体翻译质量均有显著提升。ProNMT 提供了一种高效、可扩展且具备上下文感知能力的方法,特别适用于翻译依赖上下文的表达,如代词。
原文摘要 · Abstract (English)
Pronoun translation is a longstanding challenge in neural machine translation (NMT), often requiring inter-sentential context to ensure linguistic accuracy. To address this, we introduce ProNMT, a novel framework designed to enhance pronoun and overall translation quality in context-aware machine translation systems. ProNMT leverages Quality Estimation (QE) models and a unique Pronoun Generation Likelihood-Based Feedback mechanism to iteratively fine-tune pre-trained NMT models without relying on extensive human annotations. The framework combines QE scores with pronoun-specific rewards to guide training, ensuring improved handling of linguistic nuances. Extensive experiments demonstrate significant gains in pronoun translation accuracy and general translation quality across multiple metrics. ProNMT offers an efficient, scalable, and context-aware approach to improving NMT systems, particularly in translating context-dependent elements like pronouns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。