为病原体数据设安全等级,防AI被用于生物武器
Securing Dual-Use Pathogen Data of Concern
- 提出五级生物安全数据分级框架,按风险分类病原体数据
- 每级对应不同技术限制,高风险数据需严格管控
- 适合关注AI生物安全的科研与政策制定者
训练数据是构建高性能人工智能(AI)模型的关键输入。生物AI模型依赖大量数据,包括生物序列、结构、图像和功能信息。训练数据类型直接决定模型能力,可能涉及生物安全风险。在第50届阿西洛马会议中,超过100名国际研究人员支持对数据实施管控,以防止AI被用于有害应用,如生物武器开发。为此,本文提出五级生物安全数据水平(BDL)框架,根据数据潜在贡献于生物安全关切能力的程度进行分类。每个级别包含特定数据类型,并提出相应的技术限制措施。同时,本文还设计了一套针对新生成双重用途病原体数据的新型治理框架。在计算与编程资源广泛可及的背景下,数据管控可能是降低危险生物AI能力扩散最有效的干预手段之一。
原文摘要 · Abstract (English)
Training data is an essential input into creating competent artificial intelligence (AI) models. AI models for biology are trained on large volumes of data, including data related to biological sequences, structures, images, and functions. The type of data used to train a model is intimately tied to the capabilities it ultimately possesses--including those of biosecurity concern. For this reason, an international group of more than 100 researchers at the recent 50th anniversary Asilomar Conference endorsed data controls to prevent the use of AI for harmful applications such as bioweapons development. To help design such controls, we introduce a five-tier Biosecurity Data Level (BDL) framework for categorizing pathogen data. Each level contains specific data types, based on their expected ability to contribute to capabilities of concern when used to train AI models. For each BDL tier, we propose technical restrictions appropriate to its level of risk. Finally, we outline a novel governance framework for newly created dual-use pathogen data. In a world with widely accessible computational and coding resources, data controls may be among the most high-leverage interventions available to reduce the proliferation of concerning biological AI capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。