arXiv:2508.00875cs.CYcs.AI2025-08被引 9

为通用人工智能模型评估提供可复现的全流程建议

Preliminary suggestions for rigorous GPAI model evaluations

  • 按设计、实施、执行、文档四阶段提出评估框架
  • 融合机器学习、心理学等多领域成熟方法论
  • 适合监管者、开发者及研究者参考使用

本文初步整理了通用人工智能(GPAI)评估的最佳实践,旨在提升评估的内部有效性、外部有效性和可复现性。内容涵盖人类增强研究与基准测试的建议,以及适用于多种评估类型的跨领域通用建议。建议按评估生命周期的四个阶段——设计、实施、执行和文档进行组织。借鉴机器学习、统计学、心理学、经济学、生物学等领域中已被证实有效的经验,这些实践旨在推动尚处于发展初期的GPAI评估科学领域的对话。本文目标读者包括承担系统性风险的GPAI模型提供方(欧盟人工智能法案有明确要求)、第三方评估机构、评估严谨性的政策制定者,以及开展GPAI评估的学术研究人员。

原文摘要 · Abstract (English)

This document presents a preliminary compilation of general-purpose AI (GPAI) evaluation practices that may promote internal validity, external validity and reproducibility. It includes suggestions for human uplift studies and benchmark evaluations, as well as cross-cutting suggestions that may apply to many different evaluation types. Suggestions are organised across four stages in the evaluation life cycle: design, implementation, execution and documentation. Drawing from established practices in machine learning, statistics, psychology, economics, biology and other fields recognised to have important lessons for AI evaluation, these suggestions seek to contribute to the conversation on the nascent and evolving field of the science of GPAI evaluations. The intended audience of this document includes providers of GPAI models presenting systemic risk (GPAISR), for whom the EU AI Act lays out specific evaluation requirements; third-party evaluators; policymakers assessing the rigour of evaluations; and academic researchers developing or conducting GPAI evaluations.

AI评估GPAI合规标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。