现有基准测试无法真实反映大模型的通用认知能力。
Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models
- 指出基准测试存在固有缺陷,难以衡量模型真实智能水平。
- 实验证明模型在复杂推理任务中表现脆弱,缺乏泛化能力。
- 适合关注模型评估方法论的研究者和开发者阅读。
大型语言模型(LLMs)在各类语言、知识和推理基准测试中持续展现出令人印象深刻的性能。这种快速进展使许多评论者认为,LLMs的通用认知能力也在迅速提升,暗示其在现实任务中日益强大。本文通过理论与实证分析,挑战这一观点。指出基准测试范式本身存在固有局限,且现有基准测试的特定缺陷,使其性能无法作为衡量模型可泛化认知能力的可靠指标。同时,对抗性测试与可解释性技术表明,LLMs在诸多语言与推理任务中缺乏稳健能力,常未能学习到支持泛化推理的表征。结论是:不应将基准测试成绩视为衡量大模型通用认知能力的可靠依据。
原文摘要 · Abstract (English)
Large language models (LLMs) regularly demonstrate new and impressive performance on a wide range of language, knowledge, and reasoning benchmarks. Such rapid progress has led many commentators to argue that LLM general cognitive capabilities have likewise rapidly improved, with the implication that such models are becoming progressively more capable on various real-world tasks. Here I summarise theoretical and empirical considerations to challenge this narrative. I argue that inherent limitations with the benchmarking paradigm, along with specific limitations of existing benchmarks, render benchmark performance highly unsuitable as a metric for generalisable competence over cognitive tasks. I also contend that alternative methods for assessing LLM capabilities, including adversarial stimuli and interpretability techniques, have shown that LLMs do not have robust competence in many language and reasoning tasks, and often fail to learn representations which facilitate generalisable inferences. I conclude that benchmark performance should not be used as a reliable indicator of general LLM cognitive capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。