建立统一评价标准,让所有AI模型评估结果可比可复用。
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

- 设计通用JSON schema,兼容不同来源的评估数据
- 汇聚2.2万+模型、2273个评测基准,覆盖31种格式
- 支持跨平台比对,适合研究者与工程师快速查证
AI评估广泛用于测试和理解技术进展,但评估方式多样导致结果难以比较。当前结果分散在排行榜、论文、博客、评测工具日志和自建仓库中,且不同框架对相同任务给出不一致分数,元数据记录不一,影响分析、跨社区协作、成本控制和复用。本文提出Every Eval Ever,首个共享的评估结果标准化架构与社区共建数据库。该架构以统一的JSON文档形式表示评估,具备源无关性,可接入评测工具与论文结果,并支持细粒度实例输出存储。贡献包括:(i) 首个由社区治理的元数据标准及配套实例级模式;(ii) 自动转换器,支持主流格式、评测工具与排行榜的格式迁移;(iii) 基于Hugging Face托管的众包数据库,目前已涵盖22,235个模型、2,273个独立评测基准和31种评估格式。
原文摘要 · Abstract (English)
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First, results are saved in incompatible formats, scattered across leaderboards, papers, blog posts, evaluation harness logs, and custom repositories. Second, results are created by different evaluation frameworks, which produce divergent scores for nominally identical evaluations and record metadata inconsistently, hindering comparison, cross-community evaluation science, cost reduction, and reuse. We introduce Every Eval Ever, the first shared schema and community-crowdsourced repository for AI evaluation results. The schema standardizes how evaluations are represented in a unified, single JSON document. It is source-agnostic by design, ingesting results from evaluation harnesses and papers alike, and optionally stores per-instance outputs for fine-grained analysis. We contribute: (i) a community-governed metadata schema with a companion instance-level schema, the first standardization effort of its kind; (ii) automatic converters from popular formats, evaluation harnesses, and leaderboards to the unified schema; and (iii) a crowdsourced community database hosted on Hugging Face, currently spanning to date 22,235 models, 2,273 unique benchmarks, and 31 evaluation formats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。