用视频模型生成可控图像,效果优于专门训练的图像模型。
Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis
- 将视频模型压缩为图像生成工具,利用其长程上下文能力
- 在多种控制生成任务中超越专用图像模型
- 适合想复用视频模型资源的研究者
视频生成模型因其捕捉真实世界动态连续变化的能力,可视为世界模拟器。它们整合了视觉、时间、空间和因果等高维信息,能预测不同状态下的主体。本文提出一种名为『维度缩减攻击』(DRA-Ctrl)的视频到图像知识压缩与任务适配范式,利用视频模型在长程上下文建模和全注意力机制上的优势,实现多种生成任务。针对视频帧连续性与图像生成离散性的差异,引入基于mixup的过渡策略以实现平滑适配;同时设计定制化掩码机制重构建模,更好对齐文本提示与图像级控制。在主题驱动和空间条件生成等多样任务上,经重构的视频模型表现优于直接训练于图像的数据集。结果表明,大规模视频生成器在更广泛视觉应用中具有未被发掘的潜力。DRA-Ctrl为复用资源密集型视频模型提供了新思路,并为未来跨模态统一生成模型奠定基础。
原文摘要 · Abstract (English)
Video generative models can be regarded as world simulators due to their ability to capture dynamic, continuous changes inherent in real-world environments. These models integrate high-dimensional information across visual, temporal, spatial, and causal dimensions, enabling predictions of subjects in various status. A natural and valuable research direction is to explore whether a fully trained video generative model in high-dimensional space can effectively support lower-dimensional tasks such as controllable image generation. In this work, we propose a paradigm for video-to-image knowledge compression and task adaptation, termed \textit{Dimension-Reduction Attack} (\texttt{DRA-Ctrl}), which utilizes the strengths of video models, including long-range context modeling and flatten full-attention, to perform various generation tasks. Specially, to address the challenging gap between continuous video frames and discrete image generation, we introduce a mixup-based transition strategy that ensures smooth adaptation. Moreover, we redesign the attention structure with a tailored masking mechanism to better align text prompts with image-level control. Experiments across diverse image generation tasks, such as subject-driven and spatially conditioned generation, show that repurposed video models outperform those trained directly on images. These results highlight the untapped potential of large-scale video generators for broader visual applications. \texttt{DRA-Ctrl} provides new insights into reusing resource-intensive video models and lays foundation for future unified generative models across visual modalities. The project page is https://dra-ctrl-2025.github.io/DRA-Ctrl/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。