arXiv:2509.09706cs.CRcs.AI2025-09被引 1

对比三类大模型在对抗文本攻击下的鲁棒性,发现罗伯塔和弗兰- t5 极强抗性。

Differential Robustness in Transformer Language Models: Empirical Evaluation Under Adversarial Text Attacks

  • 用 TextFooler 和 BERTAttack 设计系统攻击测试
  • 罗伯塔和弗兰-t5 攻击成功率0%,而BERT降至3%
  • 揭示当前防御机制需高算力,建议优化效率

本研究评估了大型语言模型(LLMs)在对抗性攻击下的鲁棒性,重点关注 Flan-T5、BERT 与 RoBERTa-Base。通过 TextFooler 与 BERTAttack 进行系统化对抗测试,发现模型鲁棒性差异显著:RoBERTa-Base 与 Flan-T5 表现出极强抗性,攻击成功率为 0%;而 BERT-Base 显著脆弱,TextFooler 将其准确率从 48% 降至 3%,攻击成功率达 93.75%。研究揭示部分模型虽具备有效防御机制,但通常依赖大量计算资源。本文有助于理解 LLM 安全性,识别现有防护策略的优劣,并提出更高效防御方法的实践建议。

原文摘要 · Abstract (English)

This study evaluates the resilience of large language models (LLMs) against adversarial attacks, specifically focusing on Flan-T5, BERT, and RoBERTa-Base. Using systematically designed adversarial tests through TextFooler and BERTAttack, we found significant variations in model robustness. RoBERTa-Base and FlanT5 demonstrated remarkable resilience, maintaining accuracy even when subjected to sophisticated attacks, with attack success rates of 0%. In contrast. BERT-Base showed considerable vulnerability, with TextFooler achieving a 93.75% success rate in reducing model accuracy from 48% to just 3%. Our research reveals that while certain LLMs have developed effective defensive mechanisms, these safeguards often require substantial computational resources. This study contributes to the understanding of LLM security by identifying existing strengths and weaknesses in current safeguarding approaches and proposes practical recommendations for developing more efficient and effective defensive strategies.

对抗攻击语言模型鲁棒性安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。