arXiv:2507.11968cs.CV2025-07ICCV被引 2

提出三模态攻击框架,测试短视频内容审核模型漏洞

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation

  • 设计三模态对抗攻击,同时干扰视觉、音频和语义理解
  • 在多个先进模型上实现高成功率攻击,最高达87.3%
  • 适合内容安全研究者和AI防御开发者参考

多模态大语言模型(MLLM)被广泛用于内容审核,但其在短视频场景下的鲁棒性仍待深入探索。现有安全评估多依赖单模态攻击,无法揭示联合攻击风险。本文提出全面的三模态安全评估框架:首先构建了包含多样化短视频的SVMA对抗数据集,由人工引导生成合成攻击样本;其次提出ChimeraBreak三模态攻击策略,同步挑战视觉、听觉与语义推理路径。在主流MLLM上的大量实验显示显著脆弱性,攻击成功率(ASR)普遍超过80%。分析揭示模型存在偏差,常将正常内容误判为违规,或反之。采用LLM-as-a-judge评估攻击有效性。本研究的数据集与发现为构建更鲁棒、更安全的MLLM提供了关键洞见。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are increasingly used for content moderation, yet their robustness in short-form video contexts remains underexplored. Current safety evaluations often rely on unimodal attacks, failing to address combined attack vulnerabilities. In this paper, we introduce a comprehensive framework for evaluating the tri-modal safety of MLLMs. First, we present the Short-Video Multimodal Adversarial (SVMA) dataset, comprising diverse short-form videos with human-guided synthetic adversarial attacks. Second, we propose ChimeraBreak, a novel tri-modal attack strategy that simultaneously challenges visual, auditory, and semantic reasoning pathways. Extensive experiments on state-of-the-art MLLMs reveal significant vulnerabilities with high Attack Success Rates (ASR). Our findings uncover distinct failure modes, showing model biases toward misclassifying benign or policy-violating content. We assess results using LLM-as-a-judge, demonstrating attack reasoning efficacy. Our dataset and findings provide crucial insights for developing more robust and safe MLLMs.

内容安全对抗攻击多模态视频审核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。