提出可认证的毒性感知距离,增强视觉语言模型抗攻击能力
CeTAD: Towards Certified Toxicity-Aware Distance in Vision Language Models
- 设计新距离度量捕捉恶意与正常回复的语义差异
- 通过随机平滑实现对对抗与结构扰动的正式鲁棒性保证
- 在特征空间注入噪声,防御像素级与结构级攻击
大型视觉语言模型(VLMs)在视觉理解任务中表现卓越,但其对越狱攻击的鲁棒性仍是开放挑战。本文提出一种通用可认证防御框架,严格防护VLMs免受潜在视觉越狱攻击。首先,我们提出一种新型距离度量,量化恶意响应与预期响应之间的语义差异,捕捉传统余弦相似度常忽略的细微差别。其次,设计了一种回归式认证方法,利用随机平滑技术,在黑盒设置下仍能对对抗性和结构性扰动提供形式化鲁棒性保证。此外,我们的特征空间防御在隐向量中引入噪声分布(如高斯、拉普拉斯),以抵御像素级和结构级扰动。实验表明,这种形式化且集成的策略具有构建更稳健、可信VLMs的巨大潜力。
原文摘要 · Abstract (English)
Recent advances in large vision-language models (VLMs) have demonstrated remarkable success across a wide range of visual understanding tasks. However, the robustness of these models against jailbreak attacks remains an open challenge. In this work, we propose a universal certified defence framework to safeguard VLMs rigorously against potential visual jailbreak attacks. First, we proposed a novel distance metric to quantify semantic discrepancies between malicious and intended responses, capturing subtle differences often overlooked by conventional cosine similarity-based measures. Then, we devise a regressed certification approach that employs randomized smoothing to provide formal robustness guarantees against both adversarial and structural perturbations, even under black-box settings. Complementing this, our feature-space defence introduces noise distributions (e.g., Gaussian, Laplacian) into the latent embeddings to safeguard against both pixel-level and structure-level perturbations. Our results highlight the potential of a formally grounded, integrated strategy toward building more resilient and trustworthy VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。