让流形生成模型实现精准语义编辑,可单独修改属性不扰其他。
FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers
- 利用变换器层的隐式表征构建可解释语义空间。
- 支持细粒度图像编辑与艺术创作,保持其他属性不变。
- 适用于Flux等流形生成模型,通用性强且效果稳定。
修正流模型已成为图像生成的主流方法,展现出卓越的高质量图像合成能力。然而,尽管在视觉生成上表现优异,修正流模型在图像的解耦编辑方面仍存在不足,难以实现精确的、仅影响特定属性的修改而不干扰图像其他部分。本文提出FluxSpace,一种不依赖具体领域的图像编辑方法,通过利用修正流变换器(如Flux)中变换器块所学得的表征,构建一个具备语义控制能力的表示空间。该方法生成一系列语义可解释的表示,支持从细粒度编辑到艺术创作的多种任务。本工作提供了一种可扩展且高效的图像编辑方案,具备良好的解耦能力。
原文摘要 · Abstract (English)
Rectified flow models have emerged as a dominant approach in image generation, showcasing impressive capabilities in high-quality image synthesis. However, despite their effectiveness in visual generation, rectified flow models often struggle with disentangled editing of images. This limitation prevents the ability to perform precise, attribute-specific modifications without affecting unrelated aspects of the image. In this paper, we introduce FluxSpace, a domain-agnostic image editing method leveraging a representation space with the ability to control the semantics of images generated by rectified flow transformers, such as Flux. By leveraging the representations learned by the transformer blocks within the rectified flow models, we propose a set of semantically interpretable representations that enable a wide range of image editing tasks, from fine-grained image editing to artistic creation. This work offers a scalable and effective image editing approach, along with its disentanglement capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。