arXiv:2607.21151cs.AI2026-07

视频大模型对恶意视频的拒绝能力被文本误导,导致安全失效。

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

论文配图:V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure
图 1 · 摘自论文原文
  • 通过三层次诊断框架分析模型理解与拒识的耦合问题。
  • 81%准确识别恶意视频,但恶意视频配普通问题攻击成功率仍达48.33%。
  • 视觉理解引发的拒识弱于文本理解,适合安全研究者参考。

随着视频大语言模型在真实场景中的广泛应用,其安全对齐愈发关键。反直觉的是,当恶意视频搭配良性问题时,攻击成功率高于搭配明确恶意问题的情况。为揭示该漏洞机制,本文提出V-DEAL,一个三层诊断框架,从模型行为、理解能力及内部表征层面联合分析此失败现象。通过逐步排除感知错误并量化模型内部拒识倾向,V-DEAL提供了新视角。我们在三个公开基准上测试六种视频LLM,发现模型对恶意视频识别准确率超81%,但在恶意视频配良性问题条件下,平均攻击成功率达48.33%。隐藏层分析显示,视觉理解激活的拒识倾向弱于文本理解。此外,我们提出一种提示注入干预方法,平均降低攻击成功率48.24个百分点,性能媲美以往微调方法,为视频LLM安全风险提供有效且实用的解决方案。

原文摘要 · Abstract (English)

As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuitively, we find that harmful videos paired with benign queries achieve higher attack success rates than the same videos paired with explicitly harmful queries. To understand the underlying mechanism of this vulnerability, we present V-DEAL, a three-level diagnostic framework that jointly analyzes this failure across model behaviour, understanding, and internal representations. By progressively ruling out perception failure and quantifying the model's internal refusal tendency, V-DEAL provides a new diagnostic perspective for analyzing the underlying mechanism of the observed vulnerability. We tested six Video LLMs on three public benchmarks and observed that models correctly recognize harmful video content with over 81\% accuracy, yet the average attack success rate still reaches 48.33\% under the condition pairing harmful videos with benign queries. Hidden-state analysis further shows that visual understanding activates a weaker refusal tendency than textual understanding. Furthermore, we introduce a prompt injection intervention method that reduces attack success rates by an average of 48.24 percentage points and achieves performance comparable to prior fine-tuning-based methods, providing an effective and practical means to address such safety risks in Video LLMs.

视频安全大模型对齐拒绝机制提示注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。