arXiv:2508.15361cs.CL2025-08综述被引 45

系统梳理283个大模型评测基准,揭示评估盲区与改进方向。

A Survey on Large Language Model Benchmarks

  • 按通用、领域、目标三类分类283个评测基准
  • 指出数据污染导致分数虚高、文化偏见影响公平性
  • 适合关注大模型评估体系的开发者与研究者

近年来,随着大语言模型能力的深度与广度迅速发展,各类对应的评估基准不断涌现。作为衡量模型性能的量化工具,基准不仅是评估模型能力的核心手段,也是引导模型发展方向、推动技术创新的关键要素。本文首次系统回顾大语言模型评测基准的现状与发展,将283个代表性基准分为三类:通用能力、领域特定和目标特定。通用能力基准涵盖核心语言、知识与推理;领域特定基准聚焦自然科学、人文学科与工程科技;目标特定基准关注风险、可靠性与智能体等。本文指出当前基准存在数据污染导致分数虚高、文化与语言偏见造成评估不公、缺乏对过程可信度与动态环境的评估等问题,并为未来基准设计提供可参考的范式。

原文摘要 · Abstract (English)

In recent years, with the rapid development of the depth and breadth of large language models' capabilities, various corresponding evaluation benchmarks have been emerging in increasing numbers. As a quantitative assessment tool for model performance, benchmarks are not only a core means to measure model capabilities but also a key element in guiding the direction of model development and promoting technological innovation. We systematically review the current status and development of large language model benchmarks for the first time, categorizing 283 representative benchmarks into three categories: general capabilities, domain-specific, and target-specific. General capability benchmarks cover aspects such as core linguistics, knowledge, and reasoning; domain-specific benchmarks focus on fields like natural sciences, humanities and social sciences, and engineering technology; target-specific benchmarks pay attention to risks, reliability, agents, etc. We point out that current benchmarks have problems such as inflated scores caused by data contamination, unfair evaluation due to cultural and linguistic biases, and lack of evaluation on process credibility and dynamic environments, and provide a referable design paradigm for future benchmark innovation.

大模型评估评测基准综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。