探究DNA语言模型在分类任务中的抗攻击能力,发现其极易受干扰且可提升鲁棒性。
Exploring Adversarial Robustness in Classification tasks using DNA Language Models
- 从碱基、密码子到序列层级设计对抗攻击策略
- 模型在攻击下性能显著下降,准确率大幅滑坡
- 对抗训练能同时提升鲁棒性与分类精度,适合生物信息学应用
DNA语言模型(如GROVER、DNABERT2和Nucleotide Transformer)处理的DNA序列常含测序错误、突变及实验噪声,可能严重影响模型表现。然而,此类模型的鲁棒性尚未被充分研究。本文系统评估其在分类任务中的抗扰性,采用碱基替换、密码子修改及基于反向翻译的序列变换等多层级对抗攻击策略,分析模型脆弱性。结果表明,这些模型对对抗攻击高度敏感,导致性能显著下降。进一步探索对抗训练作为防御手段,发现其可有效提升模型鲁棒性与分类准确率。本研究揭示了当前DNA语言模型的局限性,强调了在生物信息学中构建鲁棒模型的必要性。
原文摘要 · Abstract (English)
DNA Language Models, such as GROVER, DNABERT2 and the Nucleotide Transformer, operate on DNA sequences that inherently contain sequencing errors, mutations, and laboratory-induced noise, which may significantly impact model performance. Despite the importance of this issue, the robustness of DNA language models remains largely underexplored. In this paper, we comprehensivly investigate their robustness in DNA classification by applying various adversarial attack strategies: the character (nucleotide substitutions), word (codon modifications), and sentence levels (back-translation-based transformations) to systematically analyze model vulnerabilities. Our results demonstrate that DNA language models are highly susceptible to adversarial attacks, leading to significant performance degradation. Furthermore, we explore adversarial training method as a defense mechanism, which enhances both robustness and classification accuracy. This study highlights the limitations of DNA language models and underscores the necessity of robustness in bioinformatics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。