arXiv:2512.06674cs.CV2025-12被引 6

首个可自进化攻击图像到视频模型的框架,揭示其安全漏洞。

RunawayEvil: Jailbreaking the Image-to-Video Generative Models

  • 采用策略-战术-行动框架,通过强化学习和大模型实现攻击自演化。
  • 在COCO2017上攻击成功率比现有方法高出58.5%至79%。
  • 适合研究多模态安全与防御的学者,助力构建更鲁棒的视频生成系统。

图像到视频(I2V)生成从图像和文本输入中合成动态视觉内容,提供了强大的创作控制力。然而,这类多模态系统的安全性,尤其是其对越狱攻击的脆弱性,仍严重缺乏研究。为此,我们提出RunawayEvil,首个具备动态演化能力的I2V模型多模态越狱框架。该框架基于“策略-战术-行动”范式,包含三个核心组件:(1) 策略感知指令单元,通过强化学习驱动的策略定制与大模型策略探索实现攻击策略自我演化;(2) 多模态战术规划单元,根据选定策略生成协调的文本越狱指令与图像篡改指南;(3) 战术执行单元,负责执行并评估多模态协同攻击。该自演化架构使框架能在无人干预下持续适应并强化攻击策略。大量实验表明,RunawayEvil在商用I2V模型(如Open-Sora 2.0和CogVideoX)上达到当前最优攻击成功率,在COCO2017数据集上较现有方法提升58.5%至79%。本工作为I2V模型漏洞分析提供关键工具,为构建更鲁棒的视频生成系统奠定基础。

原文摘要 · Abstract (English)

Image-to-Video (I2V) generation synthesizes dynamic visual content from image and text inputs, providing significant creative control. However, the security of such multimodal systems, particularly their vulnerability to jailbreak attacks, remains critically underexplored. To bridge this gap, we propose RunawayEvil, the first multimodal jailbreak framework for I2V models with dynamic evolutionary capability. Built on a "Strategy-Tactic-Action" paradigm, our framework exhibits self-amplifying attack through three core components: (1) Strategy-Aware Command Unit that enables the attack to self-evolve its strategies through reinforcement learning-driven strategy customization and LLM-based strategy exploration; (2) Multimodal Tactical Planning Unit that generates coordinated text jailbreak instructions and image tampering guidelines based on the selected strategies; (3) Tactical Action Unit that executes and evaluates the multimodal coordinated attacks. This self-evolving architecture allows the framework to continuously adapt and intensify its attack strategies without human intervention. Extensive experiments demonstrate RunawayEvil achieves state-of-the-art attack success rates on commercial I2V models, such as Open-Sora 2.0 and CogVideoX. Specifically, RunawayEvil outperforms existing methods by 58.5 to 79 percent on COCO2017. This work provides a critical tool for vulnerability analysis of I2V models, thereby laying a foundation for more robust video generation systems.

多模态安全越狱攻击视频生成自演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。