用智能体自动构建可持续更新的评测基准,解决人工成本高、模型性能饱和问题。
Benchmark Everything Everywhere All at Once

- 部署自主智能体全程自动化生成评测任务与数据
- 自动生成15个跨领域高质量基准,人类评估表现达标
- 适合关注评测体系创新与模型能力边界的研究者
评测是衡量和推动大语言模型与多模态大模型发展的基础,提供标准化的性能度量。然而,评测构建耗时费力且难以复用,存在可持续性与可扩展性问题。此外,现有评测发布后很快达到性能饱和,难以区分先进模型。为此,我们提出Benchmark Agent——一个全自主的智能体系统,用于自动化构建评测基准。该框架可统筹从用户查询分析、子任务设计到数据标注与质量控制的完整流程。为验证其效果,我们基于该系统生成了15个代表性评测基准,涵盖文本理解、多模态理解及领域特定推理等多种场景。通过人类评估、大模型判别与一致性检查等多维度实验,证明Benchmark Agent可在极低人工干预下生成高质量评测样本。更重要的是,持续评估发现当前模型在某些领域推理任务上仍存短板。我们认为,快速迭代的评测体系将显著助力研究社区发展。演示页面与代码库将公开提供。
原文摘要 · Abstract (English)
Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance. However, their construction is labor-intensive and hard to reuse, raising concerns about sustainability and scalability. Moreover, existing benchmarks often quickly reach performance saturation after their release, resulting in insufficient discrimination among state-of-the-art models. To address these challenges, we introduce Benchmark Agent, a fully autonomous agentic system designed for benchmark building. Our framework orchestrates the complete benchmark construction pipeline, from user query analysis and subtask design to data annotation and quality control. To assess Benchmark Agent, we implement it to produce 15 representative benchmarks, spanning diverse evaluation scenarios, including text understanding, multimodal understanding, and domain-specific reasoning. Extensive experiments, including human evaluation, LLM-as-a-judge assessment, and consistency checks, demonstrate Benchmark Agent can generate high-quality benchmark samples with minimal human involvement. More importantly, through continual evaluation, we observe several insightful findings, including that current models struggle with certain domain-specific reasoning tasks. We believe that rapidly evolving benchmarks can contribute significantly to the research community. The preview and code will be publicly available at the demo page and code repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。