arXiv:2505.16643cs.CVcs.AI2025-05中稿 · ICLR被引 1

视频大模型安全漏洞严重,新框架通过双重机制提升防御能力。

From Evaluation to Defense: Advancing Safety in Video Large Language Models

  • 构建11.4万条视频-问题对的基准测试,覆盖19类风险。
  • 在多个数据集上实现最高71.1%的安全性能提升。
  • 适合关注多模态安全、模型防御的研究者与开发者。

尽管图像类大语言模型的安全风险已得到广泛研究,其视频类对应模型(Video LLMs)却仍被严重忽视。为此,我们提出VideoSafetyEval——一个大规模、真实世界视频大模型安全基准,包含11.4k个视频-查询对,覆盖19个主要风险类别。基于此,我们发现引入视频模态使安全性能平均下降34.2%,暴露出多模态攻击利用的系统性风险。为应对这一漏洞,我们提出VideoSafety-R1,一种双阶段框架,通过三项创新实现前所未有的安全提升:(1) 构建包含46k个视频-查询-思考响应三元组的VideoSafetyThinking数据集;(2) Alarm Token-Guided Safety Fine-Tuning(AT-SFT)在视觉与文本序列中注入可学习的警报标记,通过多任务目标实现跨模态的显式危害感知;(3) 安全引导的GRPO通过基于双模态验证的规则奖励动态优化策略,强化防御性推理。上述组件协同作用,推动安全对齐从危害感知转向主动推理。该框架在VSE-HH上实现71.1%的性能提升,在MMBench、VLGuard和FigStep上分别提升59.1%、44.3%和15.0%。代码与数据集已在https://github.com/Emiya-syw/VideoSafety-R1.git公开。注意:本文含有害语言与图像示例,建议谨慎阅读。

原文摘要 · Abstract (English)

While the safety risks of image-based large language models (Image LLMs) have been extensively studied, their video-based counterparts (Video LLMs) remain critically under-examined. To systematically study this problem, we introduce VideoSafetyEval - a large-scale, real-world benchmark for Video LLM safety, which comprises 11.4k video-query pairs and spans 19 principal risk categories. Based on this, we reveal that integrating video modality degrades safety performance by an average of 34.2%, thereby exposing systemic risks in multimodal attack exploitation. To address this vulnerability, we propose VideoSafety-R1, a dual-stage framework achieving unprecedented safety gains through three innovations: (1) the VideoSafetyThinking dataset contains 46k video-query-thinking response triplets; (2) Alarm Token-Guided Safety Fine-Tuning (AT-SFT) injects learnable alarm tokens into visual and textual sequences, enabling explicit harm perception across modalities via multitask objectives; and (3) safety-guided GRPO enhances defensive reasoning through dynamic policy optimization with rule-based rewards derived from dual-modality verification. These components synergize to shift safety alignment from harm perception to active reasoning. The framework achieves a 71.1% improvement on VSE-HH, and improves by 59.1%, 44.3%, and 15.0% on the image safety datasets MMBench, VLGuard, and FigStep, respectively. Our code and dataset are available at https://github.com/Emiya-syw/VideoSafety-R1.git. Note: This paper contains harmful language and image examples, and reader discretion is recommended.

视频安全多模态防御大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。