用分类器优化语义嵌入,无需提示词即可精准编辑图像。
Instructing Text-to-Image Diffusion Models via Classifier-Guided Semantic Optimization
- 通过属性分类器学习数据级语义嵌入,实现无提示词编辑。
- 在多个数据域上实现高解耦性与强泛化能力,性能优于传统提示方法。
- 无需训练或微调扩散模型,适合快速图像编辑场景。
文本到图像的扩散模型已成为高质量图像生成与编辑的强大工具。现有方法多依赖文本提示作为编辑引导,但需手动构造提示,耗时且易引入无关信息,严重限制编辑效果。本文提出一种基于属性分类器的语义嵌入优化方法,无需文本提示,也无需对扩散模型进行训练或微调,即可引导模型实现目标编辑。我们利用分类器在数据集层面学习精确的语义嵌入,其理论证明为属性语义的最优表示,从而实现解耦且准确的编辑。实验表明,该方法在不同数据领域均表现出高解耦性和强泛化能力。
原文摘要 · Abstract (English)
Text-to-image diffusion models have emerged as powerful tools for high-quality image generation and editing. Many existing approaches rely on text prompts as editing guidance. However, these methods are constrained by the need for manual prompt crafting, which can be time-consuming, introduce irrelevant details, and significantly limit editing performance. In this work, we propose optimizing semantic embeddings guided by attribute classifiers to steer text-to-image models toward desired edits, without relying on text prompts or requiring any training or fine-tuning of the diffusion model. We utilize classifiers to learn precise semantic embeddings at the dataset level. The learned embeddings are theoretically justified as the optimal representation of attribute semantics, enabling disentangled and accurate edits. Experiments further demonstrate that our method achieves high levels of disentanglement and strong generalization across different domains of data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。