用质量评估信息指导翻译纠错,减少无意义修改。
Giving the Old a Fresh Spin: Quality Estimation-Assisted Constrained Decoding for Automatic Post-Editing
- 解码时引入词级质量评估,约束纠错范围。
- 英德、英印、英马拉地语对分别提升TER 0.65、1.86、1.44点。
- 不依赖模型架构,适合各类自动校对系统使用。
自动校对(APE)系统常因过度纠错而修改本无需调整的译文,违背最小修改原则。本文提出一种新方法,在解码过程中融入词级质量评估(QE)信息,以缓解此问题。该方法与模型架构无关,可适配任意APE系统。在英德、英印、英马拉地语三组语言对上的实验表明,相比基线系统,本方法分别取得0.65、1.86、1.44点的TER改进。结果表明,QE与APE任务具有互补性,融合QE信息能有效降低APE系统的过度纠错现象。
原文摘要 · Abstract (English)
Automatic Post-Editing (APE) systems often struggle with over-correction, where unnecessary modifications are made to a translation, diverging from the principle of minimal editing. In this paper, we propose a novel technique to mitigate over-correction by incorporating word-level Quality Estimation (QE) information during the decoding process. This method is architecture-agnostic, making it adaptable to any APE system, regardless of the underlying model or training approach. Our experiments on English-German, English-Hindi, and English-Marathi language pairs show the proposed approach yields significant improvements over their corresponding baseline APE systems, with TER gains of $0.65$, $1.86$, and $1.44$ points, respectively. These results underscore the complementary relationship between QE and APE tasks and highlight the effectiveness of integrating QE information to reduce over-correction in APE systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。