用有害病毒数据微调基因语言模型,发现可恢复其滥用能力。
Open-weight genome language model safeguards: Assessing robustness via adversarial fine-tuning
- 用110种致病病毒序列微调Evo 2模型,测试安全防护有效性。
- 微调后模型在未知病毒序列上困惑度降低,预测能力显著提升。
- 即使未接触新冠序列,也能识别免疫逃逸变异株,适合研究者警惕。
新型深度学习架构正被广泛应用于生物数据,包括基因序列。这类模型称为基因语言模型(gLMs),展现出强大的预测与生成能力,引发对其可能被滥用于生成致病病毒基因组的担忧。当前主流风险缓解措施是过滤预训练数据(即从训练集中移除病毒基因组),以限制gLM在病毒相关任务上的表现。然而,尚不清楚该方法对可被敏感病原体数据微调的开源模型是否有效。本文评估了最先进的gLM Evo 2,并使用110种危害性人类感染病毒序列进行微调,以检验其滥用相关能力是否被“恢复”。结果表明,相较于预训练模型和仅用噬菌体序列微调的版本,该模型在未见病毒序列上的困惑度显著降低。此外,该模型在未接触任何SARS-CoV-2序列的情况下,仍能识别其免疫逃逸变体,达到0.6的AUROC。本研究揭示数据排除策略可能被微调手段绕过,强调亟需建立gLMs的安全框架,并提出未来在评估与缓解措施方面需进一步工作,以实现gLMs的安全部署。
原文摘要 · Abstract (English)
Novel deep learning architectures are increasingly being applied to biological data, including genetic sequences. These models, referred to as genomic language models (gLMs), have demonstrated impressive predictive and generative capabilities, raising concerns that such models may also enable misuse, for instance via the generation of genomes for human-infecting viruses. These concerns have catalyzed calls for risk mitigation measures. The de facto mitigation of choice is filtering of pretraining data (i.e., removing viral genomic sequences from training datasets) in order to limit gLM performance on virus-related tasks. However, it is not currently known how robust this approach is for securing open-source models that can be fine-tuned using sensitive pathogen data. Here, we evaluate a state-of-the-art gLM, Evo 2, and perform fine-tuning using sequences from 110 harmful human-infecting viruses to assess the rescue of misuse-relevant predictive capabilities. The fine-tuned model exhibited reduced perplexity on unseen viral sequences relative to 1) the pretrained model and 2) a version fine-tuned on bacteriophage sequences. The model fine-tuned on human-infecting viruses also identified immune escape variants from SARS-CoV-2 (achieving an AUROC of 0.6), despite having no exposure to SARS-CoV-2 sequences during fine-tuning. This work demonstrates that data exclusion might be circumvented by fine-tuning approaches that can, to some degree, rescue misuse-relevant capabilities of gLMs. We highlight the need for safety frameworks for gLMs and outline further work needed on evaluations and mitigation measures to enable the safe deployment of gLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。