对比中英文高考试题,评估大模型知识推理能力差异。
Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal Contamination
- 构建双语高考试题数据集,人工翻译确保语言质量。
- 小模型在西语中表现比英语差37%,大模型差距几乎消失。
- 结果与MMLU高度一致,小数据集可有效评估模型性能。
本文提出UNED-ACCESS 2024,一个包含1003道大学入学水平多选题的双语数据集,题目原为西班牙语,经人工翻译成英语,且未公开发布。在统一零样本设置下,对开源与专有模型进行评估,覆盖UNED-ACCESS 2024及等量MMLU子集。结果显示:(i) 推理类题目对模型仍具挑战;(ii) 小模型表现劣于大模型,且在西班牙语中退化更明显;(iii) 最优模型在双语间性能差距可忽略,小模型差距最高达37%。模型在两个语言上的排名几乎一致,且与MMLU排名具有0.98的皮尔逊相关性,表明该小规模数据集在学科维度上足够多样且具代表性。
原文摘要 · Abstract (English)
In this article we present UNED-ACCESS 2024, a bilingual dataset that consists of 1003 multiple-choice questions of university entrance level exams in Spanish and English. Questions are originally formulated in Spanish and translated manually into English, and have not ever been publicly released. A selection of current open-source and proprietary models are evaluated in a uniform zero-shot experimental setting both on the UNED-ACCESS 2024 dataset and on an equivalent subset of MMLU questions. Results show that (i) reasoning questions are challenging for models, (ii) smaller models perform worse than larger models and degrade faster in Spanish than in English and (iii) the performance gap between languages is negligible for the best models and grows up to 37% for smaller models. Model ranking on UNED-ACCESS 2024 is almost identical in English and Spanish, and has also a high correlation (0.98 Pearson) with ranking on MMLU, suggesting that a small dataset is sufficiently diverse and representative to measure performance by discipline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。