arXiv:2605.28315cs.CL2026-05

针对中英翻译知识密集领域,构建了更难的评测基准以区分模型真实能力。

HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains

论文配图:HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains
图 1 · 摘自论文原文
  • 基于12个领域、2万条双语测试项,按难度融合规则筛选
  • 跨系统得分差距扩大至FLORES-200的两倍以上,排名明显变动
  • 专为金融、医疗等专业领域设计,适合评估模型知识理解力

通用机器翻译评测基准如FLORES-200在中英双语对上已进入饱和状态,现代大语言模型得分集中于狭窄区间:22个系统在FLORES-200 zh-en GEMBA上的得分仅相差7.87分,标准差为2.29,压缩了知识密集型领域(如金融、医疗、法律、科技)中模型间的差异。为此,本文提出HardMTBench,一个面向双向中英跨领域翻译的难度感知诊断基准。该基准涵盖12个领域,包含10,000条人工精校源句与参考译文,共形成20,000个方向性测试项。通过三阶段构建流程:先生成84,566对候选数据,再用大模型多信号判别器评估知识密度、翻译难度、术语负载与参考正确性,最终依据硬度融合规则和各领域配额确定测试集。在22种系统(含通用LLM、商用引擎、专用MT模型)上,HardMTBench将系统间GEMBA得分范围扩大约两倍,引发显著排名重排,并暴露出质量指标难以捕捉的领域术语与知识缺陷。所有数据与代码已在https://github.com/jasonNLP/HardMTBench 开源。

原文摘要 · Abstract (English)

General-purpose machine translation benchmarks such as FLORES-200 have reached a saturation regime on Chinese-English pairs, where modern large language models cluster within a narrow band of high scores. Across 22 systems, FLORES-200 zh-en GEMBA scores fall in a 7.87-point range with a standard deviation of 2.29, which compresses the separation between systems on knowledge-intensive domains such as finance, healthcare, law, and science and technology. We introduce HardMTBench, a difficulty-aware diagnostic benchmark for bidirectional Chinese-English domain translation. HardMTBench covers 12 domains and contains 10,000 hand-curated source sentences with reference translations, packaged as 20,000 directional test items. A three-stage construction pipeline builds a domain-balanced candidate pool of 84{,}566 pairs, applies an LLM-based multi-signal judge over knowledge density, translation difficulty, terminology load and reference correctness, and assembles the final test set under a hardness fusion rule with per-domain quotas. Across 22 systems spanning general LLMs, commercial engines and specialised MT models, HardMTBench widens the cross-system GEMBA range by roughly a factor of two over FLORES-200, induces visible rank reorderings, and exposes domain-specific terminology and knowledge weaknesses that quality-only metrics tend to flatten. All data and code are open-sourced at https://github.com/jasonNLP/HardMTBench.

机器翻译评测基准知识密集中英翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。