测试大模型在问题改写下的表现,发现基准评估可能高估其实用性。
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
- 对六个基准题库的题目进行多种改写,测试模型鲁棒性。
- 34个模型在改写后得分显著下降,平均降幅超15%。
- 建议构建更贴近真实场景的鲁棒性评测体系。
大语言模型(LLMs)的效能通常通过MMLU、ARC-C或HellaSwag等基准测试来评估,这些问题以原始表述呈现,格式固定。然而真实应用中存在语言多样性,要求模型能应对同一问题的多种表达方式。本研究系统生成六项常见基准测试中所有问题的多种改写版本,评估34个先进大模型在不同改写输入下的表现差异。结果显示,尽管模型排名相对稳定,但绝对得分普遍下降,平均降幅超过15%。这表明大模型在语言变体面前表现脆弱,其泛化能力存疑,且当前基准评估可能无法可靠反映模型在真实场景中的实际性能。该结果质疑了现有评测方法的可靠性,强调需要发展面向鲁棒性的新型评测标准,以更真实地模拟部署环境。
原文摘要 · Abstract (English)
Large Language Models (LLMs) effectiveness is usually evaluated by means of benchmarks such as MMLU, ARC-C, or HellaSwag, where questions are presented in their original wording, thus in a fixed, standardized format. However, real-world applications involve linguistic variability, requiring models to maintain their effectiveness across diverse rewordings of the same question or query. In this study, we systematically assess the robustness of LLMs to paraphrased benchmark questions and investigate whether benchmark-based evaluations provide a reliable measure of model capabilities. We systematically generate various paraphrases of all the questions across six different common benchmarks, and measure the resulting variations in effectiveness of 34 state-of-the-art LLMs, of different size and effectiveness. Our findings reveal that while LLM rankings remain relatively stable across paraphrased inputs, absolute effectiveness scores change, and decline significantly. This suggests that LLMs struggle with linguistic variability, raising concerns about their generalization abilities and evaluation methodologies. Furthermore, the observed performance drop challenges the reliability of benchmark-based evaluations, indicating that high benchmark scores may not fully capture a model's robustness to real-world input variations. We discuss the implications of these findings for LLM evaluation methodologies, emphasizing the need for robustness-aware benchmarks that better reflect practical deployment scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。