用统计模型提升多语言评测效率与准确性,自动发现翻译错误并识别文化特异性题目。
Extending Item Response Theory for Efficient and Meaningful Multilingual Evaluation

- 基于项目反应理论扩展出多语言统一模型,分离语言与内容影响
- 预测未测试组合的准确率比现有方法低11%-16%误差,翻译错误检测覆盖全部非英语语种
- 适合需要跨语言评估大模型性能的研究者,尤其关注评测质量与公平性
多语言基准对评估大语言模型跨语言能力至关重要,但存在三大问题:评估成本随语言数量线性增长,自动翻译引入错误且难以在大规模下察觉,部分题目混淆通用知识与文化特异性知识。本文提出统一统计框架Multilingual-IRT,通过每语言难度偏移、区分度拆分(分离内容与语言效应)、每语言能力残差来解决上述问题。在包含25个LLM和29种语言的MMLU-Pro-X数据集上拟合该模型,结果表明其参数可支持三项实际应用:预测未观测的(题目, LLM, 语言)组合时,二元交叉熵比最强的基于准确率的基线低11%-16%;在所有28种非英语语言中系统性揭示候选翻译错误,而准确率基线仅集中于少数语言;恢复被准确率基线忽略的文化特异性题目。
原文摘要 · Abstract (English)
Multilingual benchmarks are central to evaluating large language models (LLMs) across languages, but they suffer from three issues: exhaustive evaluation scales linearly with the number of languages, automatic translation introduces errors that are easily missed at scale, and some items conflate general and culture-specific knowledge. We address all three with a unified statistical framework, Multilingual-IRT, which extends Item Response Theory with per-language difficulty deviations, split discriminability separating content from language effects, and per-language ability residuals. Fitting Multilingual-IRT on 25 LLMs across 29 languages of MMLU-Pro-X, we show that its fitted parameters support three practical applications: predicting unobserved (item, LLM, language) instances with 11-16% lower binary cross-entropy than the strongest accuracy-based baseline, surfacing candidate translation errors distributed across all 28 non-English languages, whereas accuracy-based baselines concentrate detections in a few languages, and recovering culture-specific items that accuracy-based baselines miss.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。