arXiv:2603.13435cs.CVcs.AI2026-03被引 4

提出新型攻击方法CtrlAttack,破坏扩散模型生成视频时的状态演化过程。

CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models

  • 用低维速度场建模扰动,通过时间积分构建连续位移场。
  • 白盒攻击成功率超90%,黑盒超80%,同时保持图像质量稳定。
  • 首次揭示视频生成模型在状态动态层面的潜在安全风险。

基于扩散的图像到视频(I2V)模型逐渐展现出类似世界模型的特性,隐式捕捉时间动态。然而,现有研究多关注视觉质量和可控性,对模型所学习的状态转移鲁棒性关注不足。为此,我们首次分析I2V模型的脆弱性,发现时间控制机制构成新的攻击面,并揭示在不同攻击设置下统一建模的挑战。基于此,提出一种轨迹控制攻击——CtrlAttack,干扰生成过程中的状态演化。具体地,将扰动表示为低维速度场,通过时间积分构建连续位移场,从而影响模型状态转移并保持时间一致性;同时将扰动映射至观测空间,使方法适用于白盒与黑盒攻击场景。实验表明,即使在低维且强正则化约束下,本方法仍能显著破坏时间一致性:白盒攻击成功率达90%以上,黑盒超80%,而FID与FVD变化分别控制在6和130以内,揭示了I2V模型在状态动力学层面的潜在安全风险。

原文摘要 · Abstract (English)

Diffusion-based image-to-video (I2V) models increasingly exhibit world-model-like properties by implicitly capturing temporal dynamics. However, existing studies have mainly focused on visual quality and controllability, and the robustness of the state transition learned by the model remains understudied. To fill this gap, we are the first to analyze the vulnerability of I2V models, find that temporal control mechanisms constitute a new attack surface, and reveal the challenge of modeling them uniformly under different attack settings. Based on this, we propose a trajectory-control attack, called CtrlAttack, to interfere with state evolution during the generation process. Specifically, we represent the perturbation as a low-dimensional velocity field and construct a continuous displacement field via temporal integration, thereby affecting the model's state transitions while maintaining temporal consistency; meanwhile, we map the perturbation to the observation space, making the method applicable to both white-box and black-box attack settings. Experimental results show that even under low-dimensional and strongly regularized perturbation constraints, our method can still significantly disrupt temporal consistency by increasing the attack success rate (ASR) to over 90% in the white-box setting and over 80% in the black-box setting, while keeping the variation of the FID and FVD within 6 and 130, respectively, thus revealing the potential security risk of I2V models at the level of state dynamics.

视频生成扩散模型安全攻击状态演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。