arXiv:2508.09662cs.CL2025-08被引 6

用少量数据高效评估大模型,兼顾公平与泛化能力。

EffiEval: Efficient and Generalizable Model Evaluation via Capability Coverage Maximization

  • 基于模型能力覆盖率最大化,自适应筛选代表性样本。
  • 仅用少量数据即可保持与全集评估一致的排名结果。
  • 无需训练,适配多模型、多数据集,适合资源受限场景。

大语言模型(LLM)的快速发展和评估基准的日益复杂,带来了巨大的计算挑战。本文提出EffiEval,一种无需训练的高效基准评估方法,在减少数据冗余的同时保持高评估可靠性。该方法满足三个核心标准:代表性(全面覆盖模型能力)、公平性(采样不依赖模型性能以避免偏差)和泛化性(可灵活迁移至不同数据集与模型族,无需大规模评估数据)。与传统依赖绝对性能或需大量评估数据的方法不同,EffiEval通过模型效用指数(MUI)自适应选择高质量代表性子集。在多个公开基准和多样化LLM上的实验表明,EffiEval仅使用原始数据的一小部分,即可实现与全集评估高度一致的排名表现。此外,方法具备良好可扩展性,用户可根据需求平衡评估效率与代表性。总体而言,EffiEval为大模型时代提供了可靠、公平且高效的评估解决方案。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) and the development of increasingly large and diverse evaluation benchmarks have introduced substantial computational challenges for model assessment. In this paper, we present EffiEval, a training-free approach for efficient benchmarking that effectively addresses data redundancy while maintaining high evaluation reliability. Our method is specifically designed to meet three key criteria for high-quality evaluation: representativeness, by ensuring comprehensive coverage of model capabilities; fairness, by remaining independent of model performance during sample selection to avoid bias; and generalizability, by enabling flexible transfer across datasets and model families without reliance on large-scale evaluation data. Unlike traditional methods that rely on absolute performance or require extensive evaluation data, our approach adaptively selects high-quality representative subsets based on the Model Utility Index (MUI). Extensive experiments on multiple public benchmarks and diverse LLMs demonstrate that EffiEval achieves strong ranking consistency with full-dataset evaluation using only a small fraction of the original data. Furthermore, our method is flexible and scalable in size, allowing users to balance evaluation efficiency and representativeness according to specific needs. Overall, EffiEval provides a practical and generalizable solution for reliable, fair, and efficient evaluation in the era of LLMs.

大模型评估高效评估代表性采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。