arXiv:2606.19899cs.CYcs.AI2026-06被引 4

评估AI在生物研究中的能力与风险,关键在于设计和解释方式。

Measuring Biological Capabilities and Risks of AI Agents

  • 通过实证评估框架分析AI在生物任务中的表现
  • 指出评估设计直接影响结果对风险的反映程度
  • 为政策制定者和科研机构提供可操作的评估指南

本文针对快速兴起的政策挑战:如何生成并解读关于AI科学家(即能自主或协作完成多步骤科学任务的代理型AI系统)的生物能力与风险的可信证据。随着这些系统进入真实科研流程,决策者面临的结果常依赖于未明示或文档不全的设计选择。本文综合现有关于AI赋能生物风险的证据,提出生物代理评估作为一种有前景但解释敏感的评估工具。核心贡献是基于实际评估经验的实践性考量,揭示定义、设计、执行、评分和记录等环节的选择如何实质性地影响评估结果对风险的含义。分析旨在帮助政策制定者谨慎解读生物评估输出;引导公私基金向高杠杆的AI-生物学评估研究投入资源;支持生物安全从业者评估新兴AI系统。次要受众包括前沿AI实验室中设计或执行代理评估的研究人员、AI提供商、科研机构及第三方评估组织。

原文摘要 · Abstract (English)

This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or collaboratively performing multi-step scientific tasks. As these systems enter real research workflows, decision-makers increasingly face evaluation results whose meaning depends on underlying design choices that are often implicit or under-documented. We synthesize current evidence on AI-enabled biological risks and introduce biological agentic evaluations as a promising, but interpretation-sensitive, tool for assessing these systems. Our central contribution is a set of practical, experience-grounded considerations -- drawing from our own evaluations -- that show how choices around defining, designing, running, scoring, and documenting evaluations materially shape what results do and do not imply about risk. The analysis is intended to help policymakers interpret biological evaluation outputs with appropriate caution; guide public and private funders toward high-leverage investments in AI-biology evaluation research; and support biosecurity practitioners assessing emerging AI systems. A secondary audience includes researchers designing or conducting agentic evaluations within frontier AI labs, AI providers, scientific institutions, and third-party evaluation organizations.

AI风险生物安全评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。