arXiv:2601.00223cs.CLcs.AI2026-01

用锚点对比较法,精准评估日英翻译模型的细微优劣。

JP-TL-Bench: Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation

  • 采用固定锚点集进行成对大模型评分,提升评估可靠性。
  • 通过布拉德利-特里模型计算胜率与0-10分的LT评分。
  • 适合迭代优化日英翻译系统,尤其关注语体与自然度差异。

我们提出JP-TL-Bench,一个轻量、开源的基准测试,用于指导日英翻译系统的迭代开发。传统评估常面临‘这两句好译文哪个更优?’而非‘这句译文是否可接受?’的问题。在日英翻译中,礼貌程度、隐含意义、省略和语域等细微差别显著影响自然度。JP-TL-Bench采用参考无依赖的成对大模型比较协议,将候选模型与固定版本化的锚点集进行对比。结果通过布拉德利-特里模型聚合,输出胜率及由逻辑变换拟合对数强度生成的标准化0-10分LT评分。由于所有候选均基于同一冻结锚点集评分,只要基础集、评判模型和聚合代码一致,得分即具结构性稳定性。

原文摘要 · Abstract (English)

We introduce JP-TL-Bench, a lightweight, open benchmark designed to guide the iterative development of Japanese-English translation systems. In this context, the challenge is often "which of these two good translations is better?" rather than "is this translation acceptable?" This distinction matters for Japanese-English, where subtle choices in politeness, implicature, ellipsis, and register strongly affect perceived naturalness. JP-TL-Bench uses a protocol built to make LLM judging both reliable and affordable: it evaluates a candidate model via reference-free, pairwise LLM comparisons against a fixed, versioned anchor set. Pairwise results are aggregated with a Bradley-Terry model and reported as win rates plus a normalized 0-10 "LT" score derived from a logistic transform of fitted log-strengths. Because each candidate is scored against the same frozen anchor set, scores are structurally stable given the same base set, judge, and aggregation code.

机器翻译评估基准日英翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。