arXiv:2605.28843cs.DLcs.CY2026-05

首次系统检测开放科研平台中的生物安全风险内容

The Biosecurity Blind Spot: Systematic Dual-use Detection in Open Science Infrastructure

论文配图:The Biosecurity Blind Spot: Systematic Dual-use Detection in Open Science Infrastructure
图 1 · 摘自论文原文
  • 用词筛选与大模型结合,分析5.2万篇预印本的潜在风险
  • 近半数研究存在双用途风险,即使初衷为公共健康
  • 建议在不牺牲透明度前提下加强元数据监管

人工智能正以前所未有的速度推动生命科学进步,涵盖蛋白质结构预测、基因组建模和药物研发(Jumper等,2021;Mak等,2024)。然而,技术快速发展与开放科学运动相结合,带来了显著的双用途研究风险,但此类问题尚未得到充分实证研究。本文首次对开放预印本平台进行系统性双用途研究关切(DURC)内容分析。我们采用混合管道对约5.2万篇2024-2025年bioRxiv预印本进行了筛查,通过词汇过滤与大语言模型评估,对标题和摘要在九类DURC、三类PEPP及五类治理类别上进行评分,参照美国和澳大利亚集团监督框架。结果显示,具有潜在双用途特征的知识内容频繁出现在公开可获取的标题与摘要中,即便研究目标为公共健康,仍常超过既定风险阈值。该映射仅反映表层信息传播,未衡量实际操作能力、下游滥用潜力或制约有害应用的技术与生物安全壁垒。我们认为,机构审查流程、资助要求及预印本平台政策需演进,以实现无损科学透明性的主动元数据监控。最终,将高风险方法的受控访问机制与科学贡献的开放摘要相结合,是规模化治理AI加速生物学的务实路径。

原文摘要 · Abstract (English)

AI is transforming life sciences research at unprecedented speed, accelerating discovery across protein structure prediction, genome modeling, and drug development (Jumper et al., 2021; Mak et al., 2024). Yet this rapid advancement, coupled with the open science movement, introduces significant dual-use research concerns that have received limited empirical scrutiny. Here we present the first systematic analysis of dual-use research of concern (DURC) content on open preprint servers. We screened ~52,000 bioRxiv preprints (2024-2025) using a hybrid pipeline of lexical filtering and large language model (LLM) evaluation, scoring metadata across nine DURC, three PEPP, and five governance categories aligned with U.S. and Australia Group oversight frameworks. Our analysis reveals that dual-use-adjacent knowledge is routinely present in openly accessible titles and abstracts, often exceeding established risk thresholds even in studies with legitimate public health objectives. While this mapping captures surface-level information diffusion, it does not measure operational capability, downstream misuse potential, or the substantial technical and biosafety barriers that constrain harmful application. We argue that institutional review processes, funding requirements, and preprint platform policies must evolve to incorporate proactive, metadata-level monitoring without compromising scientific transparency. Ultimately, harmonizing controlled-access mechanisms for high-risk methodologies with open summaries of scientific contributions offers a pragmatic framework for governing AI-accelerated biology at scale.

生物安全双用途研究预印本AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。