arXiv:2512.17146cs.CRcs.LG2025-12

用智能体框架发现基因大模型对软提示攻击的脆弱性

Biosecurity-Aware AI: Agentic Risk Auditing of Soft Prompt Attacks on ESM-Based Variant Predictors

  • 设计智能体系统自动注入软提示扰动并监测模型表现
  • 发现EPM2等顶尖基因模型在攻击下性能显著下降
  • 适合关注生物安全与基因模型可靠性的研究者

基因基础模型(GFMs),如进化规模建模(ESM),在变异效应预测中表现出色。然而,其在对抗性操纵下的安全性和鲁棒性尚未被充分探索。为此,我们提出安全智能体基因组评估器(SAGE),一个用于审计GFMs对抗脆弱性的智能体框架。SAGE通过可解释且自动化的风险审计循环运行:注入软提示扰动,监控训练检查点上的模型行为,计算AUROC和AUPR等风险指标,并生成基于大语言模型的叙述式报告。该智能体过程可在不修改底层模型的前提下持续评估嵌入空间的鲁棒性。利用SAGE,我们发现即使是先进的GFMs如ESM2也对目标软提示攻击敏感,导致可测量的性能下降。这些发现揭示了基因基础模型中此前未被察觉的关键漏洞,凸显了在临床变异解读等生物医学应用中进行智能体风险审计的重要性。

原文摘要 · Abstract (English)

Genomic Foundation Models (GFMs), such as Evolutionary Scale Modeling (ESM), have demonstrated remarkable success in variant effect prediction. However, their security and robustness under adversarial manipulation remain largely unexplored. To address this gap, we introduce the Secure Agentic Genomic Evaluator (SAGE), an agentic framework for auditing the adversarial vulnerabilities of GFMs. SAGE functions through an interpretable and automated risk auditing loop. It injects soft prompt perturbations, monitors model behavior across training checkpoints, computes risk metrics such as AUROC and AUPR, and generates structured reports with large language model-based narrative explanations. This agentic process enables continuous evaluation of embedding-space robustness without modifying the underlying model. Using SAGE, we find that even state-of-the-art GFMs like ESM2 are sensitive to targeted soft prompt attacks, resulting in measurable performance degradation. These findings reveal critical and previously hidden vulnerabilities in genomic foundation models, showing the importance of agentic risk auditing in securing biomedical applications such as clinical variant interpretation.

基因模型对抗攻击智能体审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。