arXiv:2412.16254cs.CRcs.CL2024-12中稿 · version of our pap…被引 5

动态调整模型组合,提升大模型抗攻击能力

Adversarial Robustness through Dynamic Ensemble Learning

  • 根据输入特征和攻击模式动态调整多个大模型的加权组合
  • 在多种攻击下成功率降低,准确率保持更高
  • 适合需要高安全性的实际自然语言处理场景

对抗性攻击严重威胁GPT、BERT、RoBERTa、T5等预训练语言模型的可靠性。本文提出动态集成学习增强对抗鲁棒性(ARDER),通过利用多个语言模型的多样性,并根据输入特性与检测到的对抗模式动态调整集成配置。核心组件包括用于动态加权的元模型、对抗模式检测模块,以及带有正则化的对抗训练。在标准数据集和多种对抗攻击场景下的全面评估表明,相比现有方法,ARDER显著提升了鲁棒性。通过动态重构集成以优先选择最稳健的模型,有效降低了攻击成功率,同时在对抗条件下维持更高准确率。本工作有助于构建更安全可信的AI系统,为真实NLP应用提供可落地、可扩展的对抗韧性增强方案。

原文摘要 · Abstract (English)

Adversarial attacks pose a significant threat to the reliability of pre-trained language models (PLMs) such as GPT, BERT, RoBERTa, and T5. This paper presents Adversarial Robustness through Dynamic Ensemble Learning (ARDEL), a novel scheme designed to enhance the robustness of PLMs against such attacks. ARDEL leverages the diversity of multiple PLMs and dynamically adjusts the ensemble configuration based on input characteristics and detected adversarial patterns. Key components of ARDEL include a meta-model for dynamic weighting, an adversarial pattern detection module, and adversarial training with regularization techniques. Comprehensive evaluations using standardized datasets and various adversarial attack scenarios demonstrate that ARDEL significantly improves robustness compared to existing methods. By dynamically reconfiguring the ensemble to prioritize the most robust models for each input, ARDEL effectively reduces attack success rates and maintains higher accuracy under adversarial conditions. This work contributes to the broader goal of developing more secure and trustworthy AI systems for real-world NLP applications, offering a practical and scalable solution to enhance adversarial resilience in PLMs.

对抗鲁棒性动态集成大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。