用DiT内部特征实现图像视频生成的精准控制。
ReaDiT Guidance: Control for Image and Video Generation using Diffusion Transformer Features

- 通过单个DiT模块特征实现生成过程控制
- 支持深度、姿态、边缘等空间目标引导生成
- 参数少,适用于图像与视频生成控制
我们提出DiT读出(ReaDiT)引导,一种基于扩散Transformer(DiT)模型内部特征表示的轻量级控制生成框架。ReaDiT引导利用单个DiT块的特征,根据测试时提供的空间目标(如深度图、姿态图或边缘图)引导生成过程。由于现代文生视频模型大多基于DiT主干网络,ReaDiT引导可自然扩展至视频生成,实现相机运动与动态控制。实验表明,该方法在性能上达到或优于现有基于特征和现成适配器的方法,同时所需参数更少。
原文摘要 · Abstract (English)
We present DiT Readout (ReaDiT) Guidance, a lightweight framework for controlling generation with Diffusion Transformer (DiT) models via their internal feature representations. ReaDiT Guidance uses features from a single DiT block to steer the generative process according to spatial targets - like depth, pose, or edge maps - provided at test time. Furthermore, since modern text-to-video models are largely built on DiT backbones, ReaDiT Guidance naturally extends to video generation, enabling camera and motion control. Experimental results demonstrate that our approach achieves competitive or improved results compared to existing feature-based and off-the-shelf adapter-based approaches while requiring fewer parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。