arXiv:2508.17062cs.CVcs.AI2025-08

通过空间信号引导,让视频生成更精准地遵循用户提示细节。

SSG-Dit: A Spatial Signal Guided Framework for Controllable Video Generation

  • 分两阶段生成:先用多模态模型提取空间提示,再注入扩散模型
  • 在VBench上超越现有模型,尤其在空间关系控制上表现优异
  • 轻量适配器实现高效控制,适合需要高一致性视频生成的场景

可控视频生成旨在根据用户提供的文本或初始图像精确合成视频内容。然而,现有模型常因语义一致性不足,导致生成结果偏离提示中的细微细节。为此,我们提出SSG-DiT(空间信号引导扩散Transformer),一种高效且高质量的可控视频生成框架。该方法采用解耦的两阶段流程:第一阶段通过预训练多模态模型的内部表示生成具有空间感知能力的视觉提示;第二阶段将此提示与原始文本联合作为条件,通过轻量级、参数高效的SSG-Adapter注入冻结的视频扩散Transformer主干网络。该设计采用双分支注意力机制,使模型既能利用强大的生成先验,又能被外部空间信号精确引导。大量实验表明,SSG-DiT在VBench基准测试中达到领先性能,尤其在空间关系控制和整体一致性方面显著优于现有模型。

原文摘要 · Abstract (English)

Controllable video generation aims to synthesize video content that aligns precisely with user-provided conditions, such as text descriptions and initial images. However, a significant challenge persists in this domain: existing models often struggle to maintain strong semantic consistency, frequently generating videos that deviate from the nuanced details specified in the prompts. To address this issue, we propose SSG-DiT (Spatial Signal Guided Diffusion Transformer), a novel and efficient framework for high-fidelity controllable video generation. Our approach introduces a decoupled two-stage process. The first stage, Spatial Signal Prompting, generates a spatially aware visual prompt by leveraging the rich internal representations of a pre-trained multi-modal model. This prompt, combined with the original text, forms a joint condition that is then injected into a frozen video DiT backbone via our lightweight and parameter-efficient SSG-Adapter. This unique design, featuring a dual-branch attention mechanism, allows the model to simultaneously harness its powerful generative priors while being precisely steered by external spatial signals. Extensive experiments demonstrate that SSG-DiT achieves state-of-the-art performance, outperforming existing models on multiple key metrics in the VBench benchmark, particularly in spatial relationship control and overall consistency.

视频生成扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。