自动分解扩散模型的视觉能力,用滑块控制生成效果。
SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
- 从单个文本提示中自动发现多个可解释的控制方向
- 支持组合控制,生成更丰富多样的图像变化
- 适合想探索模型潜力或进行创意设计的用户
我们提出SliderSpace,一个自动将扩散模型的视觉能力分解为可控制、易理解的方向的框架。与需手动指定每个编辑方向的传统方法不同,SliderSpace仅通过一个文本提示即可同时发现多个可解释且多样化的方向。每个方向以低秩适配器形式训练,支持组合控制,并能揭示模型隐空间中的惊喜可能性。在主流扩散模型上的大量实验表明,该方法在概念分解、艺术风格探索和多样性增强三个应用中均有效。定量评估显示,SliderSpace发现的方向能有效分解模型知识的视觉结构,揭示其隐含能力。用户研究进一步验证,相比基线方法,本方法生成的变体更具多样性且更实用。代码、数据和训练权重已公开于 https://sliderspace.baulab.info。
原文摘要 · Abstract (English)
We present SliderSpace, a framework for automatically decomposing the visual capabilities of diffusion models into controllable and human-understandable directions. Unlike existing control methods that require a user to specify attributes for each edit direction individually, SliderSpace discovers multiple interpretable and diverse directions simultaneously from a single text prompt. Each direction is trained as a low-rank adaptor, enabling compositional control and the discovery of surprising possibilities in the model's latent space. Through extensive experiments on state-of-the-art diffusion models, we demonstrate SliderSpace's effectiveness across three applications: concept decomposition, artistic style exploration, and diversity enhancement. Our quantitative evaluation shows that SliderSpace-discovered directions decompose the visual structure of model's knowledge effectively, offering insights into the latent capabilities encoded within diffusion models. User studies further validate that our method produces more diverse and useful variations compared to baselines. Our code, data and trained weights are available at https://sliderspace.baulab.info
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。