针对视频大模型的新型通用拒绝服务攻击,可大幅增加计算开销并引发安全风险。
VidDoS: Universal Denial-of-Service Attack on Video-based Large Language Models
- 采用通用优化生成无需梯度计算的干扰触发器。
- 使推理延迟增加15倍以上,生成文本长度膨胀超205倍。
- 适用于自动驾驶等实时场景,揭示严重安全隐患。
视频大模型在关键应用中日益普及,但易受能耗-延迟攻击(ELA)影响,耗尽计算资源。现有图像级方法因时间聚合机制稀释帧扰动而失效,且实时性要求使逐实例优化不切实际。本文提出首个专为视频大模型设计的通用ELA框架VidDoS。通过掩码教师强制引导模型生成高成本目标序列,结合拒绝惩罚与提前终止抑制,突破模型的简洁性偏好。在三个主流视频大模型及三个视频数据集(含视频问答与自动驾驶场景)上测试,结果表明攻击导致极端退化:生成文本长度膨胀超过205倍,推理延迟增加超过15倍。对实时自动驾驶流的模拟显示,该延迟引发致命安全违规。研究呼吁学术界重视并缓解此类高危害攻击。
原文摘要 · Abstract (English)
Video-LLMs are increasingly deployed in safety-critical applications but are vulnerable to Energy-Latency Attacks (ELAs) that exhaust computational resources. Current image-centric methods fail because temporal aggregation mechanisms dilute individual frame perturbations. Additionally, real-time demands make instance-wise optimization impractical for continuous video streams. We introduce VidDoS, which is the first universal ELA framework tailored for Video-LLMs. Our method leverages universal optimization to create instance-agnostic triggers that require no inference-time gradient calculation. We achieve this through $\textit{masked teacher forcing}$ to steer models toward expensive target sequences, combined with a $\textit{refusal penalty}$ and $\textit{early-termination suppression}$ to override conciseness priors. Testing across three mainstream Video-LLMs and three video datasets, which include video question answering and autonomous driving scenarios, shows extreme degradation. VidDoS induces a token expansion of more than 205$\times$ and inflates the inference latency by more than 15$\times$ relative to clean baselines. Simulations of real-time autonomous driving streams further reveal that this induced latency leads to critical safety violations. We urge the community to recognize and mitigate these high-hazard ELA in Video-LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。