arXiv:2609.04959cs.CL2026-09

提出新指标衡量翻译难度,关注指代依赖的远近。

Discourse Dependency: A Continuous Criterion for Translation Difficulty

论文配图:Discourse Dependency: A Continuous Criterion for Translation Difficulty
图 1 · 摘自论文原文
  • 用实体重复和代词指代定义文档内指代距离
  • 99.2%段落高指代距离时模型误判,验证其有效性
  • 发现主流数据集严重偏向低难度段落,适合评估长程依赖

当前机器翻译评测基准未明确定义难度。本文提出话语依赖(DDP),一种不依赖指标、基于源端命名实体重提及与代词共指的难度度量方法。经人工共指标注验证,99.2%的高DDP段落被正确识别为需长程上下文。将DDP应用于WMT24++与WMT25,发现两者均严重偏向低DDP段落,且领域标签无法区分。在英韩后编辑任务中,比较五种上下文注入策略,随DDP增大,所有方法均落后于人类,当DDP≥15时,人工译文显著更受青睐,但自动评估指标无差异。随着前沿系统在整体得分上趋于饱和,评估重点应从“得分高低”转向“上下文感知范围”。

原文摘要 · Abstract (English)

Recent calls for harder machine translation benchmarks have not clarified what difficulty should mean. We argue that one meaningful and currently unmeasured axis is referential reach, the distance a segment must look back into its document to resolve the entities and pronouns it contains. We formalize this as discourse dependency (DDP), a metric-free, source-side measure computed from named entity re-mentions and pronominal coreference. Validated against gold coreference, DDP errs one-sidedly in 99.2% of segments, so a high-DDP segment is certified to require long-range context. Applying DDP to WMT24++ and WMT25 shows that both are heavily skewed toward low-DDP segments, which domain labels do not distinguish. Building on DDP, we compare five context injection strategies in an English-Korean post-editing setup, varying context size and selection. As DDP grows, no strategy keeps pace with human post-editing. On segments with DDP >= 15 raters prefer human translations, while automatic metrics register no difference. As frontier systems saturate aggregate scores, DDP shifts evaluation from how well models score to how far they can reach.

翻译难度指代消解上下文建模评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。