arXiv:2506.00821cs.CRcs.AI2025-06被引 2

测试基因大模型在对抗攻击下的鲁棒性,发现其存在严重安全漏洞。

SafeGenes: Evaluating the Adversarial Robustness of Genomic Foundation Models

  • 用快速梯度符号法和软提示攻击评估基因大模型的脆弱性。
  • 蛋白BERT等模型在攻击下性能大幅下降,大模型也出现显著错误。
  • 适合关注基因组AI安全与可信性的研究人员参考。

基因组基础模型(GFMs),如进化尺度建模(ESM),在变异效应预测中表现出显著成效。然而,其对抗鲁棒性尚未得到充分探索。为此,我们提出SafeGenes:一种用于基因组基础模型安全分析的框架,利用对抗攻击评估其对精心设计的近似同源基因及嵌入空间操作的鲁棒性。本研究采用两种方法评估GFMs的对抗脆弱性:快速梯度符号法(FGSM)和软提示攻击。FGSM对输入序列施加微小扰动,而软提示攻击则通过优化连续嵌入来操纵模型预测,而不修改输入标记。结合这两种技术,SafeGenes全面评估了基因组基础模型对对抗操纵的敏感性。定向软提示攻击导致基于MLM的浅层架构(如ProteinBERT)性能严重退化,即使在高容量模型(如ESM1b和ESM1v)中也引发显著失败模式。这些发现揭示了当前基础模型中的关键安全漏洞,为提升其在高风险基因组应用(如变异效应预测)中的安全性与鲁棒性开辟了新方向。

原文摘要 · Abstract (English)

Genomic Foundation Models (GFMs), such as Evolutionary Scale Modeling (ESM), have demonstrated significant success in variant effect prediction. However, their adversarial robustness remains largely unexplored. To address this gap, we propose SafeGenes: a framework for Secure analysis of genomic foundation models, leveraging adversarial attacks to evaluate robustness against both engineered near-identical adversarial Genes and embedding-space manipulations. In this study, we assess the adversarial vulnerabilities of GFMs using two approaches: the Fast Gradient Sign Method (FGSM) and a soft prompt attack. FGSM introduces minimal perturbations to input sequences, while the soft prompt attack optimizes continuous embeddings to manipulate model predictions without modifying the input tokens. By combining these techniques, SafeGenes provides a comprehensive assessment of GFM susceptibility to adversarial manipulation. Targeted soft prompt attacks induced severe degradation in MLM-based shallow architectures such as ProteinBERT, while still producing substantial failure modes even in high-capacity foundation models such as ESM1b and ESM1v. These findings expose critical vulnerabilities in current foundation models, opening new research directions toward improving their security and robustness in high-stakes genomic applications such as variant effect prediction.

基因组AI对抗攻击模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。