arXiv:2507.00817cs.CVcs.AI2025-07被引 2

提出针对视频多模态大模型的新型对抗攻击框架,有效破坏视觉与语言融合。

CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs

  • 设计双目标语义-视觉损失函数,同时干扰文本生成和视觉表示
  • 在多个商用与开源模型上实现22.8%平均攻击提升
  • 无需显式时序约束,隐式建模时间一致性,适配多模态安全研究

视频多模态大语言模型(V-MLLMs)在时序推理与跨模态理解方面表现优异,但其对抗攻击脆弱性因复杂跨模态机制、时序依赖和计算限制而长期被忽视。本文提出CAVALRY-V(跨模态语言-视觉对抗生成框架),直接针对V-MLLM中视觉感知与语言生成的关键接口进行攻击。方法创新包括:(1) 双目标语义-视觉损失函数,同步扰乱模型的文本生成逻辑和视觉表征,破坏跨模态整合;(2) 高效两阶段生成框架,结合大规模预训练以实现跨模型迁移性,再通过专项微调保障时空一致性。在多个视频理解基准上的实证评估表明,CAVALRY-V显著优于现有攻击方法,在商业系统(GPT-4.1、Gemini 2.0)和开源模型(QwenVL-2.5、InternVL-2.5、Llava-Video、Aria、MiniCPM-o-2.6)上平均攻击效果提升22.8%。该框架通过隐式时序建模而非显式正则化实现灵活性,甚至在图像理解任务上获得34.4%的平均增益,展现出作为多模态系统对抗研究基础框架的潜力。

原文摘要 · Abstract (English)

Video Multimodal Large Language Models (V-MLLMs) have shown impressive capabilities in temporal reasoning and cross-modal understanding, yet their vulnerability to adversarial attacks remains underexplored due to unique challenges: complex cross-modal reasoning mechanisms, temporal dependencies, and computational constraints. We present CAVALRY-V (Cross-modal Language-Vision Adversarial Yielding for Videos), a novel framework that directly targets the critical interface between visual perception and language generation in V-MLLMs. Our approach introduces two key innovations: (1) a dual-objective semantic-visual loss function that simultaneously disrupts the model's text generation logits and visual representations to undermine cross-modal integration, and (2) a computationally efficient two-stage generator framework that combines large-scale pre-training for cross-model transferability with specialized fine-tuning for spatiotemporal coherence. Empirical evaluation on comprehensive video understanding benchmarks demonstrates that CAVALRY-V significantly outperforms existing attack methods, achieving 22.8% average improvement over the best baseline attacks on both commercial systems (GPT-4.1, Gemini 2.0) and open-source models (QwenVL-2.5, InternVL-2.5, Llava-Video, Aria, MiniCPM-o-2.6). Our framework achieves flexibility through implicit temporal coherence modeling rather than explicit regularization, enabling significant performance improvements even on image understanding (34.4% average gain). This capability demonstrates CAVALRY-V's potential as a foundational approach for adversarial research across multimodal systems.

对抗攻击视频理解多模态安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。