arXiv:2509.19100cs.LGcs.AI2025-09被引 2

提升深度学习模型在视觉、医疗和语言任务中的抗攻击能力

Algorithms for Adversarially Robust Deep Learning

  • 提出新训练范式与认证算法增强视觉模型鲁棒性
  • 在医学影像等任务中实现当前最优的跨域泛化性能
  • 针对大模型越狱攻击,开发前沿的攻防技术

鉴于深度学习模型在安全关键应用中的广泛应用,确保其决策对对抗性攻击具备鲁棒性至关重要。本文探讨了近年来设计具备理想鲁棒性特性的算法的进展。首先讨论计算机视觉中的对抗样本问题,提出新的技术成果、训练范式与认证算法。其次研究领域泛化问题,即训练神经网络从一组训练分布推广到未见的测试分布,提出新算法在医学影像、分子识别和图像分类任务中达到当前最优泛化性能。最后研究大型语言模型(LLM)越狱场景,即攻击者设计提示以诱导模型生成不当内容,提出新型攻击与防御方法,代表了构建鲁棒语言智能体的前沿进展。

原文摘要 · Abstract (English)

Given the widespread use of deep learning models in safety-critical applications, ensuring that the decisions of such models are robust against adversarial exploitation is of fundamental importance. In this thesis, we discuss recent progress toward designing algorithms that exhibit desirable robustness properties. First, we discuss the problem of adversarial examples in computer vision, for which we introduce new technical results, training paradigms, and certification algorithms. Next, we consider the problem of domain generalization, wherein the task is to train neural networks to generalize from a family of training distributions to unseen test distributions. We present new algorithms that achieve state-of-the-art generalization in medical imaging, molecular identification, and image classification. Finally, we study the setting of jailbreaking large language models (LLMs), wherein an adversarial user attempts to design prompts that elicit objectionable content from an LLM. We propose new attacks and defenses, which represent the frontier of progress toward designing robust language-based agents.

对抗鲁棒性领域泛化大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。