arXiv:2605.18824cs.LGcs.AI2026-05

自动生成覆盖广、带元数据的细粒度评测基准,提升大模型评估可靠性。

Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models

论文配图:Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models
图 1 · 摘自论文原文
  • 基于教科书等参考材料,多智能体生成问题并构建解题图以确保答案可信。
  • 在机器学习、企业金融和个人金融领域生成的新基准错误率显著低于MMLU和GSM8K。
  • 可揭示现有基准忽略的模型性能差异,适合需要精细评估的研究者使用。

大模型评估常依赖于覆盖不全且缺乏元数据的综合评分基准。本文提出一种自动化基准生成框架,基于教科书等参考材料生成问题,构建涵盖面广、元数据丰富且抗污染能力强的评测集。该流程采用多智能体架构生成问题,并通过解题图驱动策略显著提升真实答案的可靠性。基于此框架,我们生成了机器学习、企业金融和个人金融三个领域的评测集。专家评审显示,其真实答案错误率显著低于MMLU和GSM8K等现有基准。对12个商用与开源模型的评估表明,新基准实现了近均匀的能力覆盖,并揭示了现有基准无法捕捉的模型间性能差异。框架与精选评测集即将开源。

原文摘要 · Abstract (English)

Evaluation of foundation models often rely on aggregate scores from benchmarks that lack comprehensive coverage and metadata for a fine-grained evaluation. We introduce a framework for automated benchmark generation. Our framework generates evaluation problems grounded in reference material, such as textbooks, producing benchmarks with broad coverage, rich metadata, and robustness to contamination. The pipeline employs a multi-agent architecture for problem generation and a solution-graph-driven strategy that significantly improves the reliability of ground truth solutions. Using the framework, we generate three benchmarks in Machine Learning, Corporate Finance, and Personal Finance. Expert review finds a significantly lower ground-truth error rate than previous benchmarks such as MMLU and GSM8K. Evaluation of 12 commercial and open-source models shows that our benchmarks achieve near-uniform competency coverage and surface performance differences across models that existing benchmarks fail to capture. We will open-source the framework and our curated benchmarks soon.

评测基准大模型评估自动化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。