用少量低质合成数据就能高效实现文本到视频的可控生成。
Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation
- 用稀疏低质合成数据替代真实高清数据微调模型。
- 在控制精度和视觉质量上优于使用真实数据微调的模型。
- 适合资源有限但需快速部署可控生成的开发者。
将大规模文本到视频扩散模型微调以添加新的生成控制(如相机快门速度、光圈等)通常需要大量高质量数据,获取困难。本文提出一种数据高效的微调策略,仅需稀疏且低质量的合成数据即可学习这些控制。实验表明,使用此类简单数据进行微调不仅实现了所需控制,还取得了优于使用逼真“真实”数据微调的效果。此外,我们构建了一个框架,从直观和定量两个层面解释该现象。
原文摘要 · Abstract (English)
Fine-tuning large-scale text-to-video diffusion models to add new generative controls, such as those over physical camera parameters (e.g., shutter speed or aperture), typically requires vast, high-fidelity datasets that are difficult to acquire. In this work, we propose a data-efficient fine-tuning strategy that learns these controls from sparse, low-quality synthetic data. We show that not only does fine-tuning on such simple data enable the desired controls, it actually yields superior results to models fine-tuned on photorealistic "real" data. Beyond demonstrating these results, we provide a framework that justifies this phenomenon both intuitively and quantitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。