arXiv:2508.10175cs.CL2025-08EMNLP被引 10

自动识别机器翻译难题文本,助力更精准评估与研究方向定位。

Estimating Machine Translation Difficulty

  • 基于翻译质量预期定义文本难度,构建可量化评估的难度估计任务。
  • 新模型Sentinel-src-24/25在挑战性基准上显著优于启发式方法与大模型评分。
  • 适用于筛选高难度语料,指导模型改进与评测体系升级。

近年来机器翻译质量持续提升,在主流评测中已接近完美,导致难以区分先进模型并识别改进方向。本文提出翻译难度估计任务,将文本难度定义为翻译质量的预期水平。我们引入新评估指标,用于衡量难度估计器性能,并对比基线与新方法。实验表明,专用模型显著优于基于规则的方法和大模型作为裁判的方案,其中Sentinel-src表现最佳。为此,我们发布两个优化模型:Sentinel-src-24与Sentinel-src-25,可用于大规模文本扫描,筛选出最可能挑战当前机器翻译系统的内容,从而构建更具挑战性的评测基准。

原文摘要 · Abstract (English)

Machine translation quality has steadily improved over the years, achieving near-perfect translations in recent benchmarks. These high-quality outputs make it difficult to distinguish between state-of-the-art models and to identify areas for future improvement. In this context, automatically identifying texts where machine translation systems struggle holds promise for developing more discriminative evaluations and guiding future research. In this work, we address this gap by formalizing the task of translation difficulty estimation, defining a text's difficulty based on the expected quality of its translations. We introduce a new metric to evaluate difficulty estimators and use it to assess both baselines and novel approaches. Finally, we demonstrate the practical utility of difficulty estimators by using them to construct more challenging benchmarks for machine translation. Our results show that dedicated models outperform both heuristic-based methods and LLM-as-a-judge approaches, with Sentinel-src achieving the best performance. Thus, we release two improved models for difficulty estimation, Sentinel-src-24 and Sentinel-src-25, which can be used to scan large collections of texts and select those most likely to challenge contemporary machine translation systems.

机器翻译难度估计评测基准模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。