为AI评估报告设计可解释的统一卡片,让结果更透明可比。
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

- 构建统一报告框架,整合评测元数据与模型信息
- 覆盖5816个模型、10万+结果,发现报告系统性缺失
- 支持科研与非科研读者的不同解读需求
AI评估结果虽大规模生成,但报告方式在排行榜、模型卡片、基准论文和公司博客间不一致,导致读者难以可靠比较、识别遗漏或追溯结论依据。现有工作仅解决局部问题,仍存在三大缺口:仅覆盖评估生命周期片段且无法组合成可读记录;采用静态表示,无法适应不同利益相关方的提问需求;仅为纸上提案,缺乏规模化部署所需的提取基础设施。本文提出\EvalCards{},一个可操作的报告层,将基准元数据、评估运行数据与模型元数据融合为统一记录。我们(1)通过分析52篇论文和10次利益相关方访谈,推导出报告模式;(2)实现四种可解释信号(可复现性、文档完整性、来源与风险、得分可比性),并通过针对科研与非科研受众的阅读模式呈现;(3)部署监控工具,在5,816个模型、635个基准、101,843项结果上应用\EvalCards{},揭示当前报告实践中的系统性缺陷。
原文摘要 · Abstract (English)
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers cannot reliably compare results across sources, identify what a report omits, or trace an aggregate claim to its underlying evidence. Recent efforts address isolated components but leave three gaps: they cover only narrow slices of the evaluation lifecycle and do not compose into a single interpretable record; they specify static representations that do not differentiate the questions different stakeholders bring to the same evidence; and they remain proposals on paper, lacking the extraction infrastructure required for adoption at scale. We present \EvalCards{}, an operational reporting layer that composes benchmark metadata, evaluation run data, and model metadata into a unified record. We (1) derive a reporting schema from a structured review of 52 papers and 10 stakeholder interviews, (2) implement four interpretive signals (reproducibility, documentation completeness, provenance and risk, and score comparability), rendered through reader modes calibrated to research and non-research audiences, and (3) deploy a monitoring tool that applies \EvalCards{} across 5,816 models, 635 benchmarks, and 101,843 results, surfacing systematic gaps in current reporting practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。