arXiv:2608.25089cs.CL2026-08

跨语言模型评估需谨慎,常用指标有偏见,语义对齐的负对数似然更可靠。

Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

论文配图:Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation
图 1 · 摘自论文原文
  • 用平行语料训练可控单语模型,测试不同评估方法
  • 多种标准化指标因分词等差异引入跨语言偏差
  • 基于语义等价句子的负对数似然提供更一致的比较

跨语言语言模型评估的公平性仍是多语言自然语言处理中的基础挑战。现有研究采用多种下游任务和内在指标,理论依据各异,但极少实证检验这些方法是否产生有意义的跨语言结论。本文系统考察了使用平行数据训练的可控单语模型,通过调整分词器词汇量和模型规模,进一步在多语言大模型上验证结果。研究表明,几种广泛使用的归一化指标因分词、编码和正字法差异引入跨语言偏差。相比之下,在语义等价序列上计算的句级负对数似然能提供更合理且一致的跨语言比较。

原文摘要 · Abstract (English)

Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP. Existing studies adopt a variety of downstream tasks and intrinsic metrics with different theoretical justifications, yet there has been little empirical investigation into whether these approaches yield meaningful crosslingual conclusions. We systematically examine crosslingual evaluation approaches using controlled monolingual language models trained on parallel data with varying tokenizer vocabulary sizes and model sizes, and further validate our findings on multilingual LLMs. We further discuss challenges in achieving comparable downstream evaluation across languages. Our results show that several widely used normalized metrics introduce crosslinguistic biases rooted in tokenization, encoding, and orthographic differences. In contrast, sentence-level negative log-likelihood computed over semantically equivalent sequences provides more meaningful and consistent crosslingual comparisons.

跨语言评估语言模型偏差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。