无需反向传播即可精准控制视频生成中物体位置。
Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization

- 用解析解替代梯度优化,直接求解空间引导。
- 定位精度显著优于现有方法,计算开销几乎为零。
- 适合需要高效可控视频生成的科研与应用开发。
扩散变换器文本到视频模型虽已实现卓越合成质量,但精细的空间可控性仍是重大挑战。现有无训练方法在空间定位生成方面表现良好,却依赖代价高昂的梯度优化技术,尤其在大型模型中计算开销成倍增加。为此,我们提出一种全新的无训练、无梯度方法——梯度自由解析轨迹优化视频生成(GATO-Vid),用于精确空间引导。不同于传统反向传播,我们引入一种替代交叉注意力得分,并通过解析方式求解,获得精确闭式解。为使用该解析解,我们设计了一种针对Transformer潜在空间拓扑结构的实时注入机制。实验表明,GATO-Vid在定位准确性上显著优于现有基线,同时带来极小的计算开销。
原文摘要 · Abstract (English)
Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant challenge. While existing training-free methods produce solid overall results in spatially grounded generation, \ie, placing a specific object in a designated location, they rely on gradient-based optimization techniques that incur prohibitive computational overhead, a bottleneck amplified in modern large-scale architectures. To address this limitation, we present Gradient-free Analytical Trajectory Optimization Video Generation (GATO-Vid), a novel training-free and gradient-free approach for precise spatial guidance. Rather than relying on costly backward passes, we introduce an alternative cross-attention score and solve it analytically to obtain an exact, closed-form solution. To use our analytical solution, we propose an on-the-fly injection mechanism tailored to the topological manifold of the transformer's latent space. Our experiments demonstrate that GATO-Vid significantly outperforms existing baselines in localization accuracy while introducing minimal computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。