发现多语言评测集存在严重错误,改后模型表现差超10%
Test Set Quality in Multilingual LLM Evaluation
- 人工检查法语和泰卢固语评测集,发现大量错误
- 修正后,多个大模型在两语言上得分差异达近10%
- 建议评测集应可迭代、版本化,避免误判模型能力
近年来,为评估大语言模型的多语言能力,出现了若干半自动构建的多语言基准数据集。然而,对数据集本身质量的关注仍不足,尽管已有研究指出完全人工标注的测试集也存在错误。本文对法语和泰卢固语两个语言的最新多语言评估集进行了人工分析,发现了多项错误。对比多个大模型在原始与修正版本数据集上的表现,发现在两种语言中均出现显著差异(某些情况相差接近10%)。基于此,我们主张测试集不应视为不可更改,而应定期复查、验证并支持版本管理。最后,我们向数据集创建者与使用者提出改进建议,以应对评测集质量问题。
原文摘要 · Abstract (English)
Several multilingual benchmark datasets have been developed in a semi-automatic manner in the recent past to measure progress and understand the state-of-the-art in the multilingual capabilities of Large Language Models. However, there is not a lot of attention paid to the quality of the datasets themselves, despite the existence of previous work in identifying errors in even fully human-annotated test sets. In this paper, we manually analyze recent multilingual evaluation sets in two languages - French and Telugu, identifying several errors in the process. We compare the performance difference across several LLMs with the original and revised versions of the datasets and identify large differences (almost 10% in some cases) in both languages). Based on these results, we argue that test sets should not be considered immutable and should be revisited, checked for correctness, and potentially versioned. We end with some recommendations for both the dataset creators as well as consumers on addressing the dataset quality issues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。