发现机器翻译中的推理错误,但修正效果有限。
Should We be Pedantic About Reasoning Errors in Machine Translation?
- 用自动标注法识别翻译推理中的三类错位错误。
- 小修正对翻译质量影响小,强干预虽提升纠错率但效果不一。
- 在乌尔都语中可高精度识别错误,但修正后问题仍存。
在多个语言对(英→西、法、德、中文、日、乌尔都、粤语)中,我们发现了机器翻译中的推理错误。为量化错误发生频率,采用自动化标注协议检测推理步骤是否属于三类错误:(1)与源句错位,(2)与模型假设错位,(3)推理路径错位。通过一系列从弱到强的干预手段(如缓和、移除、重推理、事后修正、最优干预)对有误推理路径进行修复。实验表明,小幅度修正对翻译质量影响甚微,而强干预虽能显著提升纠错率,但翻译质量改善不一致。在乌尔都语中可高精度识别错误,西班牙语中精度较低;然而,即便消除推理错误,初始翻译问题也未明显缓解,表明机器翻译的推理忠实性有限。
原文摘要 · Abstract (English)
Across multiple language pairings (English $\to$ \{Spanish, French, German, Mandarin, Japanese, Urdu, Cantonese\}), we find reasoning errors in translation. To quantify how often these reasoning errors occur, we leverage an automated annotation protocol for reasoning evaluation wherein the goal is to detect if a reasoning step is any of three error categories: (1) source sentence-misaligned, (2) model hypothesis-misaligned, or (3) reasoning trace-misaligned. We probe the reasoning model with perturbed traces correcting for these identified reasoning errors using an array of weak-to-strong interventions: hedging, removal, re-reasoning after removal, hindsight, and oracle interventions. Experimenting with interventions on the reasoning traces suggests that small corrections to the reasoning have little impact on translation quality, but stronger interventions yield the highest resolution rates, despite translation quality gains being mixed. We find ultimately that reasoning errors in MT can be identified with high precision in Urdu but lower precision in Spanish, but that removing these reasoning errors does not resolve the initial errors significantly, suggesting limited reasoning faithfulness for machine translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。