让图像属性独立调控更精准,支持多属性同时滑动控制。
CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation
- 在条件先验潜空间中设计可组合滑块,解耦多属性控制
- 引入新损失函数,实现多属性调整时保持图像结构一致
- 无需重训练主模型,适配图像与视频生成任务
在文本到图像生成中,即使使用详细提示词,对年龄、微笑等属性的细粒度控制仍具挑战。滑块方法可提供精确控制,但现有方法为每个属性单独训练适配器,忽略多属性间的纠缠,导致属性间干扰,难以同时精确调控多个属性。为此,我们提出CompSlider,通过生成条件先验实现多属性的解耦控制,可在不重训练基础模型的前提下,同时调节多个属性。我们引入新的解耦与结构损失,确保多属性变化时图像结构保持一致。由于在条件先验的潜空间中操作,该方法显著降低训练与推理的计算开销。我们在多种图像属性上进行评估,并通过扩展至视频生成展示了其通用性。
原文摘要 · Abstract (English)
In text-to-image (T2I) generation, achieving fine-grained control over attributes - such as age or smile - remains challenging, even with detailed text prompts. Slider-based methods offer a solution for precise control of image attributes. Existing approaches typically train individual adapter for each attribute separately, overlooking the entanglement among multiple attributes. As a result, interference occurs among different attributes, preventing precise control of multiple attributes together. To address this challenge, we aim to disentangle multiple attributes in slider-based generation to enbale more reliable and independent attribute manipulation. Our approach, CompSlider, can generate a conditional prior for the T2I foundation model to control multiple attributes simultaneously. Furthermore, we introduce novel disentanglement and structure losses to compose multiple attribute changes while maintaining structural consistency within the image. Since CompSlider operates in the latent space of the conditional prior and does not require retraining the foundation model, it reduces the computational burden for both training and inference. We evaluate our approach on a variety of image attributes and highlight its generality by extending to video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。