arXiv:2608.08283cs.CLcs.AI2026-08

测试现有翻译评估指标在古汉英翻译中的可靠性。

Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

  • 设计最小差异对诊断古汉英翻译错误类型。
  • 所有指标均有盲区,MetricX24表现最佳。
  • 适合数字人文领域研究者参考改进评估方法。

尽管大语言模型能意外地处理部分历史语言翻译,但其在数字人文工作流中的应用受限于缺乏可靠的评估手段。本文以古汉英翻译为案例,考察现有针对现代语言设计的自动评估指标在此场景下的可靠性。我们提出一种基于最小差异对的诊断框架,捕捉学术使用中关键的错误类型,测试了基于参考和无参考的多种指标对错误的敏感性及对合理变化的容忍度。结果发现,所有指标均存在盲区,其中 MetricX24 整体表现最优。研究强调需开发更鲁棒、可解释的评估指标,以适应历史与文化差异显著的翻译场景。

原文摘要 · Abstract (English)

Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, however MetricX24 performs best overall. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings.

机器翻译评估指标数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。