大模型评估正从固定测试转向动态能力衡量,解决泛化难题。
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- 用核心能力替代具体任务构建评估体系
- 引入大模型自评与动态数据集提升效率
- 适合关注评估方法演进的研究者与工程师
大语言模型(LLMs)发展迅猛,已广泛应用于学术、产业和日常场景。为应对这一趋势,本文系统分析了大模型兴起带来的评估挑战。我们识别并探讨两大关键转变:(i) 从任务特定评估转向基于能力的评估,将基准测试重构为知识、推理、指令遵循、多模态理解及安全等核心能力维度;(ii) 从人工评估转向自动化评估,涵盖动态数据集构建与“大模型作为评判者”(LLM-as-a-judge)的评分机制。然而,即便如此,一个核心障碍依然存在:评估泛化问题。有限的测试集无法跟上模型能力近乎无限增长的步伐。本文将从方法、数据集、评估者和度量标准四个角度剖析此问题及前述转变中的关键挑战。鉴于该领域快速演进,本文维护一个可动态更新的 GitHub 仓库(各节附链接),欢迎贡献者与合作者持续完善。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are advancing at an amazing speed and have become indispensable across academia, industry, and daily applications. To keep pace with the status quo, this survey probes the core challenges that the rise of LLMs poses for evaluation. We identify and analyze two pivotal transitions: (i) from task-specific to capability-based evaluation, which reorganizes benchmarks around core competencies such as knowledge, reasoning, instruction following, multi-modal understanding, and safety; and (ii) from manual to automated evaluation, encompassing dynamic dataset curation and "LLM-as-a-judge" scoring. Yet, even with these transitions, a crucial obstacle persists: the evaluation generalization issue. Bounded test sets cannot scale alongside models whose abilities grow seemingly without limit. We will dissect this issue, along with the core challenges of the above two transitions, from the perspectives of methods, datasets, evaluators, and metrics. Due to the fast evolving of this field, we will maintain a living GitHub repository (links are in each section) to crowd-source updates and corrections, and warmly invite contributors and collaborators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。