通过低秩适配器实现图像视频生成中概念的高效连续控制
Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters
- 在预训练文本编码器中识别低秩方向,实现轻量级概念控制
- 训练速度比Concept Slider快5倍,比Attribute Control快47倍
- 无需重训练即可适配不同模型,适合需要灵活编辑的生成场景
扩散模型在图像与视频生成方面取得显著进展。尽管已有多种概念控制方法可实现对自由文本提示的细粒度、连续且灵活控制,但这些方法通常需大量训练时间与显存,且需针对不同扩散主干网络重新训练,限制了其可扩展性与适应性。为此,我们提出Text Slider,一种轻量、高效且即插即用的框架,通过在预训练文本编码器中识别低秩方向,实现视觉概念的连续控制,显著降低训练时间、显存消耗和可训练参数数量。此外,Text Slider支持多概念组合与连续调节,可在图像与视频生成中实现精细灵活的操控。实验表明,Text Slider能平滑调节特定属性,同时保持输入的原始空间布局与结构。相较而言,其训练速度比Concept Slider快5倍,比Attribute Control快47倍,显存使用分别减少近2倍和4倍。
原文摘要 · Abstract (English)
Recent advances in diffusion models have significantly improved image and video synthesis. In addition, several concept control methods have been proposed to enable fine-grained, continuous, and flexible control over free-form text prompts. However, these methods not only require intensive training time and GPU memory usage to learn the sliders or embeddings but also need to be retrained for different diffusion backbones, limiting their scalability and adaptability. To address these limitations, we introduce Text Slider, a lightweight, efficient and plug-and-play framework that identifies low-rank directions within a pre-trained text encoder, enabling continuous control of visual concepts while significantly reducing training time, GPU memory consumption, and the number of trainable parameters. Furthermore, Text Slider supports multi-concept composition and continuous control, enabling fine-grained and flexible manipulation in both image and video synthesis. We show that Text Slider enables smooth and continuous modulation of specific attributes while preserving the original spatial layout and structure of the input. Text Slider achieves significantly better efficiency: 5$\times$ faster training than Concept Slider and 47$\times$ faster than Attribute Control, while reducing GPU memory usage by nearly 2$\times$ and 4$\times$, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。