提出评估生物大模型双用途风险的框架,发现现有过滤无效
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
- 构建三维度评估体系,检测病毒序列、突变和毒力理解能力
- 实测显示过滤后数据可快速通过微调恢复,且泛化能力强
- 提示仅靠数据过滤难保安全,需更深层防护策略
开放权重的生物基础模型存在双重用途风险:虽能加速科研与药物研发,也可能被恶意利用制造致命生物武器。当前方法主要在预训练阶段过滤危险数据,但其有效性尚未明确,尤其对有动机的攻击者可通过微调实现恶意用途。为此,我们提出BioRiskEval框架,从序列建模、突变效应预测和毒力预测三个角度评估模型对病毒的理解能力。结果表明,现有过滤措施效果有限——部分被排除的知识可通过微调快速恢复,且在序列建模中展现出更强泛化性;此外,双用途信号可能已存在于预训练表征中,仅通过简单线性探测即可激发。这些发现揭示了仅依赖数据过滤的局限性,凸显了对开放权重生物模型亟需更稳健的安全与防护策略。
原文摘要 · Abstract (English)
Open-weight bio-foundation models present a dual-use dilemma. While holding great promise for accelerating scientific research and drug development, they could also enable bad actors to develop more deadly bioweapons. To mitigate the risk posed by these models, current approaches focus on filtering biohazardous data during pre-training. However, the effectiveness of such an approach remains unclear, particularly against determined actors who might fine-tune these models for malicious use. To address this gap, we propose BioRiskEval, a framework to evaluate the robustness of procedures that are intended to reduce the dual-use capabilities of bio-foundation models. BioRiskEval assesses models' virus understanding through three lenses, including sequence modeling, mutational effects prediction, and virulence prediction. Our results show that current filtering practices may not be particularly effective: Excluded knowledge can be rapidly recovered in some cases via fine-tuning, and exhibits broader generalizability in sequence modeling. Furthermore, dual-use signals may already reside in the pretrained representations, and can be elicited via simple linear probing. These findings highlight the challenges of data filtering as a standalone procedure, underscoring the need for further research into robust safety and security strategies for open-weight bio-foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。