arXiv:2504.12661cs.LGcs.CL2025-04ACL被引 11

通过推理驱动的提示优化,主动提升视觉语言模型的安全性。

VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization

  • 用多模态推理重写用户输入,动态识别图文交互中的潜在风险。
  • 在SIUO基准上,平均安全性能提升43.59%,优于四个基线方法。
  • 无需修改模型参数,适配多种VLM架构,适合部署级安全增强。

将视觉语言模型(VLM)与安全标准对齐对于缓解其多模态复杂性带来的风险至关重要,因视觉与语言融合会暴露传统防护难以察觉的细微威胁。受跨模态推理是预判复杂漏洞关键的启发,我们提出一种新的VLM安全方向:多模态推理驱动的提示重写。为此,我们引入VLMGuard-R1,一个主动式框架,通过推理引导的重写器优化用户输入,动态解析文本-图像交互,生成增强安全性的精炼提示,且不改变VLM核心参数。我们设计三阶段推理流程构建数据集,训练重写器以推断隐蔽威胁,实现定制化、可操作的响应而非泛化拒绝。在三个基准上对五种VLM进行的广泛实验表明,VLMGuard-R1优于四个基线。尤其在SIUO基准上,五种模型的平均安全性能提升43.59%。

原文摘要 · Abstract (English)

Aligning Vision-Language Models (VLMs) with safety standards is essential to mitigate risks arising from their multimodal complexity, where integrating vision and language unveils subtle threats beyond the reach of conventional safeguards. Inspired by the insight that reasoning across modalities is key to preempting intricate vulnerabilities, we propose a novel direction for VLM safety: multimodal reasoning-driven prompt rewriting. To this end, we introduce VLMGuard-R1, a proactive framework that refines user inputs through a reasoning-guided rewriter, dynamically interpreting text-image interactions to deliver refined prompts that bolster safety across diverse VLM architectures without altering their core parameters. To achieve this, we devise a three-stage reasoning pipeline to synthesize a dataset that trains the rewriter to infer subtle threats, enabling tailored, actionable responses over generic refusals. Extensive experiments across three benchmarks with five VLMs reveal that VLMGuard-R1 outperforms four baselines. In particular, VLMGuard-R1 achieves a remarkable 43.59\% increase in average safety across five models on the SIUO benchmark.

视觉语言模型安全对齐提示优化推理驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。