arXiv:2509.20792cs.CVcs.AI2025-09中稿 · ICCV被引 1

用动态对抗课程提升视觉语言模型的少样本适应鲁棒性

DAC-LoRA: Dynamic Adversarial Curriculum for Efficient and Robust Few-Shot Adaptation

  • 设计渐进式对抗攻击课程,智能提升攻击难度
  • 在保持高准确率的同时显著增强抗攻击能力
  • 可无缝接入现有轻量微调流程,适合安全关键场景

视觉语言模型(VLMs)广泛应用于自动驾驶、医疗诊断和内容审核等关键领域。尽管参数高效微调(PEFT)方法如LoRA能高效适配特定任务,但这些模型仍易受对抗攻击影响,威胁安全决策。以CLIP为骨干的VLMs是高价值目标,其漏洞可能在整个多模态AI生态中传播。本文提出动态对抗课程框架DAC-LoRA,将对抗训练融入PEFT。核心思想是基于一阶驻点条件(FOSC)与TRADES启发的损失,构建逐步升级的攻击课程,适用于任何迭代攻击方法。实验表明,该方法在不显著牺牲干净准确率的前提下,大幅提升了模型对抗鲁棒性,且实现轻量化、通用性强,可轻松集成至标准PEFT流程。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are foundational to critical applications like autonomous driving, medical diagnosis, and content moderation. While Parameter-Efficient Fine-Tuning (PEFT) methods like LoRA enable their efficient adaptation to specialized tasks, these models remain vulnerable to adversarial attacks that can compromise safety-critical decisions. CLIP, the backbone for numerous downstream VLMs, is a high-value target whose vulnerabilities can cascade across the multimodal AI ecosystem. We propose Dynamic Adversarial Curriculum DAC-LoRA, a novel framework that integrates adversarial training into PEFT. The core principle of our method i.e. an intelligent curriculum of progressively challenging attack, is general and can potentially be applied to any iterative attack method. Guided by the First-Order Stationary Condition (FOSC) and a TRADES-inspired loss, DAC-LoRA achieves substantial improvements in adversarial robustness without significantly compromising clean accuracy. Our work presents an effective, lightweight, and broadly applicable method to demonstrate that the DAC-LoRA framework can be easily integrated into a standard PEFT pipeline to significantly enhance robustness.

对抗训练少样本学习LoRA视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。