arXiv:2502.16361cs.CYcs.CL2025-02被引 3

提出新框架量化视觉语言模型安全风险,助力公共领域可信AI落地

A Framework for Evaluating Vision-Language Model Safety: Building Trust in AI for Public Sector Applications

  • 通过噪声扰动分析模型脆弱区域,识别误分类阈值
  • 构建复合噪声图与显著性模式,发现模型易受攻击部位
  • 设计综合漏洞评分,融合随机噪声与对抗攻击影响

视觉语言模型(VLMs)在公共部门应用日益广泛,亟需对其安全性与对抗攻击脆弱性进行可靠评估。本文提出一种新框架,量化VLMs的对抗风险。我们分析模型在高斯噪声、椒盐噪声和均匀噪声下的表现,识别误分类阈值,并推导出复合噪声块与显著性模式,揭示模型脆弱区域。这些模式与快速梯度符号法(FGSM)对比,评估其对抗有效性。本文还提出新的漏洞评分(Vulnerability Score),综合随机噪声与对抗攻击的影响,提供全面的模型鲁棒性评估指标。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are increasingly deployed in public sector missions, necessitating robust evaluation of their safety and vulnerability to adversarial attacks. This paper introduces a novel framework to quantify adversarial risks in VLMs. We analyze model performance under Gaussian, salt-and-pepper, and uniform noise, identifying misclassification thresholds and deriving composite noise patches and saliency patterns that highlight vulnerable regions. These patterns are compared against the Fast Gradient Sign Method (FGSM) to assess their adversarial effectiveness. We propose a new Vulnerability Score that combines the impact of random noise and adversarial attacks, providing a comprehensive metric for evaluating model robustness.

视觉语言模型模型安全对抗攻击公共应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。