arXiv:2501.02654cs.CLcs.AI2025-01被引 4

构建更严格的文本对抗防御基准,推动模型鲁棒性研究

Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks

  • 构建涵盖多任务、多数据集的综合性文本对抗防御评估框架
  • 首次系统评估主流防御方法在分类、语义匹配、常识推理等任务上的表现
  • 为研究人员提供新标准,助力提升NLP模型抗攻击能力

近期自然语言处理进展揭示了深度学习模型对对抗攻击的脆弱性。尽管已提出多种防御机制,但缺乏覆盖多样数据集、模型和任务的综合评估基准。本文提出一个全面的文本对抗防御基准,显著超越以往工作。该基准涵盖广泛数据集,评估前沿防御方法,并扩展至单句分类、语义相似性与改写识别、自然语言推理及常识推理等关键任务。本工作不仅为对抗鲁棒性领域的研究者与实践者提供宝贵资源,还指明未来研究的关键方向。通过建立该领域的新评估标准,旨在加速构建更鲁棒、可靠的自然语言处理系统。

原文摘要 · Abstract (English)

Recent advancements in natural language processing have highlighted the vulnerability of deep learning models to adversarial attacks. While various defence mechanisms have been proposed, there is a lack of comprehensive benchmarks that evaluate these defences across diverse datasets, models, and tasks. In this work, we address this gap by presenting an extensive benchmark for textual adversarial defence that significantly expands upon previous work. Our benchmark incorporates a wide range of datasets, evaluates state-of-the-art defence mechanisms, and extends the assessment to include critical tasks such as single-sentence classification, similarity and paraphrase identification, natural language inference, and commonsense reasoning. This work not only serves as a valuable resource for researchers and practitioners in the field of adversarial robustness but also identifies key areas for future research in textual adversarial defence. By establishing a new standard for benchmarking in this domain, we aim to accelerate progress towards more robust and reliable natural language processing systems.

对抗防御NLP鲁棒性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。