arXiv:2410.22584cs.LGcs.AI2024-10被引 2

用多智能体自动构建高质量评测基准,提升模型评估效率与准确性。

BenchAgents: Multi-Agent Systems for Structured Benchmark Creation

  • 通过多智能体协作分解评测构建流程,实现自动化生成。
  • 在语言与视觉任务中验证了规划、约束满足等能力的评测基准效果。
  • 适合研究者用于发现模型缺陷和评估新生成能力。

评估效果受限于高质量基准的可用性。随着模型发展,亟需针对新型复杂生成能力设计评测基准。然而,人工构建基准耗时费力,难以全面覆盖各类能力。我们提出BenchAgents,一个基于大语言模型的多智能体框架,可系统化地自动化创建评测基准,并内建保障数据与评价指标质量。该框架将基准创建过程分解为规划、生成、验证与评估四个阶段,由多个LLM智能体协同完成。智能体间交互并根据开发者反馈动态优化数据多样性与质量。我们利用BenchAgents构建了涵盖语言与视觉模态的规划、约束满足及因果推理能力的评测基准,并基于这些基准分析当前先进模型,揭示常见失败模式与模型差异。

原文摘要 · Abstract (English)

Evaluation insights are limited by the availability of high-quality benchmarks. As models evolve, there is a need to create benchmarks that can measure progress on new and complex generative capabilities. However, manually creating new benchmarks is slow and expensive, restricting comprehensive evaluations for any capability. We introduce BenchAgents, a multi-agent framework that methodically leverages large language models (LLMs) to automate evaluation benchmark creation while inherently ensuring data and (evaluation) metric quality. BenchAgents decomposes the benchmark creation process into planning, generation, verification, and evaluation, each of which is ] orchestrated via LLM agents. These agents interact with each other and utilize feedback from benchmark developers to improve and flexibly control data diversity and quality. We use BenchAgents to create benchmarks to evaluate capabilities related to planning, constraint satisfaction, and causal reasoning spanning both language and vision modalities. We then use these benchmarks to study state-of-the-art models and extract new insights into common failure modes and model differences.

多智能体评测基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。