用自动方法精准评估大模型能力,减少人工成本。
Automated Capability Evaluation of Foundation Models
- 利用前沿模型分解领域为语义能力,自动生成多样评测任务。
- 数学领域生成433个能力、1.18万任务,覆盖94%维基技能。
- 仅测半数能力即逼近全量评估效果,适合高效能力分析。
当前大模型评估依赖静态人工基准,难以全面捕捉模型能力。本文提出主动能力评估框架ACE,利用先进模型将领域分解为语义能力,并自动生成多样化评测任务,显著降低人工成本。在数学领域,ACE生成433个能力与11,800个任务,覆盖维基百科定义的94%技能,并发现新且连贯的能力。通过在潜在语义空间拟合能力模型,结合主动学习仅需评估少于一半能力,即可达到与全量评估相差0.01 RMSE的效果。相比静态数据集,ACE提供更均衡覆盖,揭示聚合指标忽略的细粒度差异。结果表明,ACE能更完整、深入地描绘模型能力,对大模型安全部署至关重要。
原文摘要 · Abstract (English)
Current evaluation frameworks for foundation models rely heavily on static, manually curated benchmarks, limiting their ability to capture the full breadth of model capabilities. This paper introduces Active learning for Capability Evaluation (ACE), a novel framework for scalable, automated, and fine-grained evaluation of foundation models. ACE leverages the knowledge embedded in powerful frontier models to decompose a domain into semantically meaningful capabilities and generates diverse evaluation tasks, significantly reducing human effort. In Mathematics, ACE generated 433 capabilities and 11,800 tasks, covering 94% of Wikipedia-defined skills in the domain while introducing novel, coherent ones. To maximize efficiency, ACE fits a capability model in latent semantic space, allowing reliable approximation of a subject model's performance by evaluating only a subset of capabilities via active learning. It reaches within 0.01 RMSE of exhaustive evaluation by evaluating less than half of capabilities. Compared to static datasets, ACE provides more balanced coverage and uncovers fine-grained differences that aggregate metrics fail to capture. Our results demonstrate that ACE provides a more complete and informative picture of model capabilities, which is essential for safe and well-informed deployment of foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。