arXiv:2505.11842cs.CVcs.CL2025-05NeurIPS被引 28

首个评估视频版大模型安全性的基准,发现视频能诱导67.2%的恶意攻击

Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs

  • 分拆视频语义为对象与运动文本,可控生成安全测试视频
  • 用良性问题配特定视频可触发67.2%的攻击成功率
  • 引入新评分机制,提升对模糊危害输出的判断准确性

大型视觉语言模型(LVLM)在实际应用中面临潜在恶意输入的安全风险。现有多模态安全评测主要针对静态图像输入,忽视了视频的时间动态特性可能带来的独特安全隐患。为此,我们提出Video-SafetyBench,首个专门用于评估视频-文本攻击下LVLM安全性的综合性基准。该基准包含2,264个视频-文本对,覆盖48个细粒度不安全类别,每对包含一个合成视频与一个有害查询(含明确恶意)或看似无害但结合视频会引发不良行为的良性查询。为生成语义准确的测试视频,我们设计了一种可控流水线,将视频语义分解为主体图像(展示什么)和运动文本(如何移动),共同指导生成与查询相关的视频。为有效评估不确定或边界情况下的有害输出,我们提出RJScore——一种基于LLM的新指标,融合判别模型置信度与人类对齐的决策阈值校准。大量实验表明,使用良性查询组合视频可达到平均67.2%的攻击成功率,揭示出对视频诱导攻击的普遍脆弱性。我们相信Video-SafetyBench将推动未来基于视频的安全评估与防御研究。

原文摘要 · Abstract (English)

The increasing deployment of Large Vision-Language Models (LVLMs) raises safety concerns under potential malicious inputs. However, existing multimodal safety evaluations primarily focus on model vulnerabilities exposed by static image inputs, ignoring the temporal dynamics of video that may induce distinct safety risks. To bridge this gap, we introduce Video-SafetyBench, the first comprehensive benchmark designed to evaluate the safety of LVLMs under video-text attacks. It comprises 2,264 video-text pairs spanning 48 fine-grained unsafe categories, each pairing a synthesized video with either a harmful query, which contains explicit malice, or a benign query, which appears harmless but triggers harmful behavior when interpreted alongside the video. To generate semantically accurate videos for safety evaluation, we design a controllable pipeline that decomposes video semantics into subject images (what is shown) and motion text (how it moves), which jointly guide the synthesis of query-relevant videos. To effectively evaluate uncertain or borderline harmful outputs, we propose RJScore, a novel LLM-based metric that incorporates the confidence of judge models and human-aligned decision threshold calibration. Extensive experiments show that benign-query video composition achieves average attack success rates of 67.2%, revealing consistent vulnerabilities to video-induced attacks. We believe Video-SafetyBench will catalyze future research into video-based safety evaluation and defense strategies.

视频安全大模型评测攻击防御多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。