arXiv:2512.04062cs.LG2025-12被引 7

为AI评估设计标准化文档框架,提升透明度与可复现性。

Eval Factsheets: A Structured Framework for Documenting AI Evaluations

  • 构建五维评估文档框架:上下文、范围、结构、方法、对齐性。
  • 通过问卷形式强制记录关键信息,确保评估可比较。
  • 适用于传统基准和大模型评分等多种评估方式,适合研究者与评审者使用。

基准测试的快速增多带来了可复现性、透明度和明智决策的挑战。尽管数据集和模型已有如数据清单和模型卡等结构化文档框架,评估方法却缺乏系统性文档标准。本文提出Eval Factsheets,一种基于综合分类法和问卷式方法的结构化描述框架,用于系统记录AI系统评估。该框架涵盖五个核心维度:上下文(谁在何时进行评估?)、范围(评估什么?)、结构(评估如何构建?)、方法(如何运作?)和对齐性(是否可靠、有效、鲁棒?)。我们将其转化为一个包含五个部分的实用问卷,含必填与推荐项。通过多个基准的案例研究,证明Eval Factsheets能有效捕捉从传统基准到大模型作为评判者的多样化评估范式,同时保持一致性和可比性。我们希望该框架被纳入现有及新发布的评估体系,推动更高水平的透明度与可复现性。

原文摘要 · Abstract (English)

The rapid proliferation of benchmarks has created significant challenges in reproducibility, transparency, and informed decision-making. However, unlike datasets and models -- which benefit from structured documentation frameworks like Datasheets and Model Cards -- evaluation methodologies lack systematic documentation standards. We introduce Eval Factsheets, a structured, descriptive framework for documenting AI system evaluations through a comprehensive taxonomy and questionnaire-based approach. Our framework organizes evaluation characteristics across five fundamental dimensions: Context (Who made the evaluation and when?), Scope (What does it evaluate?), Structure (With what the evaluation is built?), Method (How does it work?) and Alignment (In what ways is it reliable/valid/robust?). We implement this taxonomy as a practical questionnaire spanning five sections with mandatory and recommended documentation elements. Through case studies on multiple benchmarks, we demonstrate that Eval Factsheets effectively captures diverse evaluation paradigms -- from traditional benchmarks to LLM-as-judge methodologies -- while maintaining consistency and comparability. We hope Eval Factsheets are incorporated into both existing and newly released evaluation frameworks and lead to more transparency and reproducibility.

评估框架可复现性透明度AI评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。