arXiv:2508.09853cs.CYcs.AI2025-08被引 12

为AI模型评估报告制定透明标准,提升化学生物安全测试可信度

STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports

  • 提出STREAM标准,规范模型报告中化学生物评估的披露方式
  • 通过23位专家共识设计,确保评估过程可查、结果可信
  • 提供模板和示范案例,适合开发者与第三方审核者使用

对危险AI能力的评估对于管控灾难性风险至关重要。公开透明地披露评估内容——包括测试范围、实施方式及结果如何影响决策——是建立对AI发展信任的关键。本文提出STREAM(ChemBio)标准,旨在改进模型报告中评估结果的披露方式,初期聚焦化学与生物(ChemBio)基准测试。该标准在政府、民间组织、学术界及前沿AI企业共23位专家共同参与下制定,具备两大目标:一是作为实用工具,帮助AI开发者更清晰呈现评估结果;二是协助第三方判断模型报告是否包含足够细节以评估化学生物评估的严谨性。文中提供了“黄金标准”示例,并附三页可直接使用的报告模板,助力开发者落地实施。

原文摘要 · Abstract (English)

Evaluations of dangerous AI capabilities are important for managing catastrophic risks. Public transparency into these evaluations - including what they test, how they are conducted, and how their results inform decisions - is crucial for building trust in AI development. We propose STREAM (A Standard for Transparently Reporting Evaluations in AI Model Reports), a standard to improve how model reports disclose evaluation results, initially focusing on chemical and biological (ChemBio) benchmarks. Developed in consultation with 23 experts across government, civil society, academia, and frontier AI companies, this standard is designed to (1) be a practical resource to help AI developers present evaluation results more clearly, and (2) help third parties identify whether model reports provide sufficient detail to assess the rigor of the ChemBio evaluations. We concretely demonstrate our proposed best practices with "gold standard" examples, and also provide a three-page reporting template to enable AI developers to implement our recommendations more easily.

AI安全评估标准透明报告

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。