用大模型做翻译评估,能省35倍思考成本还更准。
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
- 让大模型学人类思考路径,提升翻译评估能力。
- 在WMT24上实现7B到32B模型相关性提升8.7点。
- 适合需要高精度自动翻译评估的研究者使用。
大推理模型(LRM)通过中间“思考”过程提升了复杂任务的推理能力,但其在机器翻译(MT)质量评估中的潜力尚未被充分探索。本文首次系统分析了以大模型为评判者的翻译评估效果,发现其需定制化评估材料,对简单样本易过度思考,且评分机制易导致高估。为此,我们提出通过合成的人类式思考轨迹训练来校准大模型的推理过程。在WMT24评测基准上的实验表明,该方法将思考开销降低约35倍,同时在不同规模的模型(7B至32B)上均提升评估性能,例如R1-Distill-Qwen-7B的相关性提升8.7点。结果表明,经校准的大模型可高效推动细粒度自动翻译评估的发展。
原文摘要 · Abstract (English)
Recent advancements in large reasoning models (LRMs) have introduced an intermediate "thinking" process prior to generating final answers, improving their reasoning capabilities on complex downstream tasks. However, the potential of LRMs as evaluators for machine translation (MT) quality remains underexplored. We provides the first systematic analysis of LRM-as-a-judge in MT evaluation. We identify key challenges, revealing LRMs require tailored evaluation materials, tend to "overthink" simpler instances and have issues with scoring mechanisms leading to overestimation. To address these, we propose to calibrate LRM thinking by training them on synthetic, human-like thinking trajectories. Our experiments on WMT24 Metrics benchmarks demonstrate that this approach largely reduces thinking budgets by ~35x while concurrently improving evaluation performance across different LRM scales from 7B to 32B (e.g., R1-Distill-Qwen-7B achieves a +8.7 correlation point improvement). These findings highlight the potential of efficiently calibrated LRMs to advance fine-grained automatic MT evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。