让文字生成模型增强或抑制特定概念,无需添加新元素。
Scaling Concept With Text-Guided Diffusion Models
- 通过分解文本引导扩散模型中的概念,实现对概念强度的调节。
- 在WeakConcept-10数据集上验证,可有效提升模糊概念的表现。
- 支持图像与音频领域的零样本应用,如姿态标准化与声音突出/消除。
文本引导扩散模型通过文本描述生成高质量内容,也实现了概念替换的编辑范式(如将狗替换为虎)。本文探索新方法:不替换概念,而是增强或抑制其本身。通过实证研究发现,概念可在文本引导扩散模型中被分解。基于此,提出ScalingConcept——一种简单有效的技术,可在真实输入中放大或缩小已分解概念,无需引入新元素。为系统评估该方法,构建了WeakConcept-10数据集,其中概念不完整需增强。更重要的是,ScalingConcept实现了跨图像与音频领域的多种零样本应用,如生成标准姿态和生成性声音突出或移除。
原文摘要 · Abstract (English)
Text-guided diffusion models have revolutionized generative tasks by producing high-fidelity content from text descriptions. They have also enabled an editing paradigm where concepts can be replaced through text conditioning (e.g., a dog to a tiger). In this work, we explore a novel approach: instead of replacing a concept, can we enhance or suppress the concept itself? Through an empirical study, we identify a trend where concepts can be decomposed in text-guided diffusion models. Leveraging this insight, we introduce ScalingConcept, a simple yet effective method to scale decomposed concepts up or down in real input without introducing new elements. To systematically evaluate our approach, we present the WeakConcept-10 dataset, where concepts are imperfect and need to be enhanced. More importantly, ScalingConcept enables a variety of novel zero-shot applications across image and audio domains, including tasks such as canonical pose generation and generative sound highlighting or removal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。