用大模型自动生成可靠通用的评测集,效率比人工高百倍。
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
- 构建四维十项评估框架,自动验证并优化大模型生成的评测题
- 在12个大模型上测试,评测结果相关性达0.967,接近人工水准
- 每题仅需0.38分钟、0.005美元,适合大规模评测场景
大语言模型的快速发展带来了模型数量与应用需求的激增。为实现有效匹配,亟需可靠、通用且高效的评测集生成工具。然而,人工标注效率低下,现有基于大模型的评测生成方法既缺乏泛化能力,又因缺少全面的验证与优化框架而可靠性不足。为此,我们首先提出一个自动化、无偏见的评估框架,涵盖四个维度和十个标准。在此基础上,系统分析直接提示大模型作为通用评测生成器的优劣。为提升可靠性,我们针对发现的问题设计一系列改进方法,并集成为BenchMaker。在多个大模型和任务上的实验表明,BenchMaker在所有指标上均达到或超过人工标注评测集的表现,展现出卓越的泛化性与可靠性。更重要的是,其在12个大模型上的评估结果高度一致(与MMLU-Pro的相关性为0.967),且每样本耗时仅0.38分钟,成本低至0.005美元。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has led to a surge in both model supply and application demands. To facilitate effective matching between them, reliable, generic and efficient benchmark generators are widely needed. However, human annotators are constrained by inefficiency, and current LLM benchmark generators not only lack generalizability but also struggle with limited reliability, as they lack a comprehensive evaluation framework for validation and optimization. To fill this gap, we first propose an automated and unbiased evaluation framework, structured around four dimensions and ten criteria. Under this framework, we carefully analyze the advantages and weaknesses of directly prompting LLMs as generic benchmark generators. To enhance the reliability, we introduce a series of methods to address the identified weaknesses and integrate them as BenchMaker. Experiments across multiple LLMs and tasks confirm that BenchMaker achieves superior or comparable performance to human-annotated benchmarks on all metrics, highlighting its generalizability and reliability. More importantly, it delivers highly consistent evaluation results across 12 LLMs (0.967 Pearson correlation against MMLU-Pro), while taking only $0.005 and 0.38 minutes per sample.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。