arXiv:2605.07093cs.CLcs.AI2026-05

中文多语言基准的翻译偏差并非固定值,而是依赖于评估方法和题目类型。

The Translation Tax Is Not a Scalar: A Counterfactual Audit of English-Source Cue Inheritance in Chinese Multilingual Benchmarks

  • 通过改写中文表达测试翻译残留影响,发现不同题目反应各异
  • 高残留题目受益明显,低残留题目无提升,无统一翻译税
  • 提出可复现的验证协议与报告清单,供同行审计使用

翻译税常被视为一个固定数值:认为翻译后的评测集会因保留英文源线索而虚高分数。本文在英译中场景下对此进行反事实审计。三种代理估计器结果不一:回译差异小且对解析器敏感;线索得分校准无法预测个体题目的增益;六模型原生对照显示,模型族效应远大于统一基准效应。新增同题型大模型自然化压力测试,保持答案、选项与内容不变,仅重写中文表述。修正提示构建错误后,模型族交互不再显著,但仍存残余剂量反应关系:高残留题目获益,低残留题目无变化。结果表明,翻译税并非单一数值,而是依赖评估方式与题目标识的风险集合。论文发布逐单元证据、自然化流程、人工质检数据及报告检查清单,供翻译多语言评测研究参考。

原文摘要 · Abstract (English)

The Translation Tax is often treated as a scalar: translated benchmarks are assumed to inflate scores by preserving English-source cues. We audit this claim in an English-to-Chinese setting. Three proxy estimators disagree: back-translation gaps are small and parser-fragile; cue-score calibration does not predict item-level gains; and a six-model native-control comparison shows model-family rather than uniform benchmark effects. We add a same-item LLM-naturalization stress test that holds answer, options, and content fixed while rewriting Chinese surface form. After correcting a prompt-construction bug, this contrast no longer supports a model-family interaction, but it preserves a residue dose-response: high-residue items benefit while low-residue items do not. The result is not a single Translation Tax, but a set of estimator- and item-dependent validity risks. We release per-cell evidence, the naturalization protocol, human QC, and a reporting checklist for translated multilingual benchmark papers.

多语言评测翻译偏差反事实审计可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。