arXiv:2511.21695cs.CLcs.AI2025-11被引 6

为提升NLP评估透明度,提出标准化报告卡片框架

EvalCards: A Framework for Standardized Evaluation Reporting

  • 设计EvalCards评估披露卡统一报告格式
  • 解决可复现性、可访问性与治理三方面问题
  • 适合研究者与开发者用于模型评估报告

评估一直是自然语言处理的核心议题,如今在大量开放模型快速发布背景下,透明的报告实践比以往更加关键。通过对近期评估与文档工作的调研,我们识别出当前报告实践中的三个长期问题:可复现性、可访问性和治理。我们认为现有标准化努力仍显不足,提出评估披露卡(EvalCards)作为改进路径。EvalCards旨在增强研究人员和实践者之间的透明度,同时为应对日益增长的治理要求提供实用基础。

原文摘要 · Abstract (English)

Evaluation has long been a central concern in NLP, and transparent reporting practices are more critical than ever in today's landscape of rapidly released open-access models. Drawing on a survey of recent work on evaluation and documentation, we identify three persistent shortcomings in current reporting practices: reproducibility, accessibility, and governance. We argue that existing standardization efforts remain insufficient and introduce Evaluation Disclosure Cards (EvalCards) as a path forward. EvalCards are designed to enhance transparency for both researchers and practitioners while providing a practical foundation to meet emerging governance requirements.

评估标准透明报告NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。