通过频率分解实现图像主体与风格解耦,提升定制化图像的精准度。
Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image Customization

- 利用图像频域特性分离主体内容与风格特征,独立优化嵌入
- 在多个数据集上主体保真度和文本匹配度均超越主流方法
- 适合需要高精度图像定制且关注风格独立控制的场景
图像定制化旨在从参考图像中学习目标主体,并根据文本提示生成条件图像,主要修改风格或背景。现有方法通常通过微调将多种概念属性整合到统一潜在嵌入中,但属性纠缠导致难以消除风格与背景中的无关干扰。为此,我们提出等衡扩散(Equilibrated Diffusion),一种基于频率驱动的方法,通过解耦纠缠的概念特征,实现均衡定制与一致的文本-视觉匹配。不同于传统方法使用共享嵌入统一微调,本工作利用图像频域成分与语义之间的内在关联:低频代表主体内容,高频对应风格。我们在频域空间分解概念并独立优化各嵌入。这种分离优化使去噪器能够捕捉脱离主体身份的风格,更好地泛化至未见风格提示。多频嵌入融合保留了模型原有的空间定制能力。我们进一步引入掩码引导扩散以限制无关背景变化并增强文本对齐。空间注意力中插入残差参考注意力(RRA),以保持主体结构与身份一致性。实验表明,等衡扩散在主体保真度和文本遵循性方面均优于主流基线,验证了方法的优势。
原文摘要 · Abstract (English)
Image customization learns target subjects from reference concept images and generates conditioned images per text prompts, mainly modifying styles or backgrounds. Prevailing methods adopt fine-tuning to pack diverse concept attributes into a unified latent embedding, yet entangled attributes hinder elimination of irrelevant disturbances from style and background. To address this issue, we propose Equilibrated Diffusion, a frequency-driven approach that disentangles tangled concept features for balanced customization and consistent text-visual matching. Unlike conventional methods learning full concepts with shared embeddings and unified tuning, our work utilizes the inherent link between image frequency components and semantics: low frequencies represent subject content and high frequencies correspond to styles. We decompose concepts in frequency space and optimize each embedding independently. This separate optimization enables the denoiser to capture style detached from subject identity and generalize better to unseen stylistic prompts. Merging multi-frequency embeddings preserves the model's original spatial customization ability. We further deploy mask-guided diffusion to restrict irrelevant background changes and boost text alignment. Residual Reference Attention (RRA) is inserted into spatial attention to retain subject structure and identity consistency. Experiments prove Equilibrated Diffusion exceeds mainstream baselines on subject fidelity and text adherence, verifying our method's superiority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。