arXiv:2504.17269cs.CV2025-04被引 1

无需训练即可实现多种语义编辑,且适配各类扩散模型。

Towards Generalized and Training-Free Text-Guided Semantic Manipulation

  • 通过控制噪声几何关系实现语义编辑,无需微调或优化。
  • 支持添加、删除、风格迁移等多种操作,结果保持高保真。
  • 插件式设计,兼容多模态任务,通用性强适合快速部署。

文本引导的语义编辑指将由源提示生成的图像修改为匹配目标提示,实现期望的语义变化(如增加、删除、风格迁移),同时保留无关内容。基于扩散模型的强大生成能力,该任务展现出生成高质量视觉内容的潜力。然而,现有方法通常需要耗时微调(效率低)、难以支持多种语义操作(可扩展性差),或不支持不同模态任务(泛化性有限)。深入分析发现,扩散模型中噪声的几何特性与语义变化强相关。受此启发,我们提出一种新型的GTF方法,具备两大优势:1)通用性:支持多种语义操作(如添加、删除、风格迁移),可无缝集成至所有基于扩散模型的方法中(即插即用),且跨模态通用;2)免训练:仅通过控制噪声间的几何关系即可生成高质量结果,无需任何调优或优化。大量实验验证了该方法的有效性,展现了其推动语义编辑技术发展的潜力。

原文摘要 · Abstract (English)

Text-guided semantic manipulation refers to semantically editing an image generated from a source prompt to match a target prompt, enabling the desired semantic changes (e.g., addition, removal, and style transfer) while preserving irrelevant contents. With the powerful generative capabilities of the diffusion model, the task has shown the potential to generate high-fidelity visual content. Nevertheless, existing methods either typically require time-consuming fine-tuning (inefficient), fail to accomplish multiple semantic manipulations (poorly extensible), and/or lack support for different modality tasks (limited generalizability). Upon further investigation, we find that the geometric properties of noises in the diffusion model are strongly correlated with the semantic changes. Motivated by this, we propose a novel $\textit{GTF}$ for text-guided semantic manipulation, which has the following attractive capabilities: 1) $\textbf{Generalized}$: our $\textit{GTF}$ supports multiple semantic manipulations (e.g., addition, removal, and style transfer) and can be seamlessly integrated into all diffusion-based methods (i.e., Plug-and-play) across different modalities (i.e., modality-agnostic); and 2) $\textbf{Training-free}$: $\textit{GTF}$ produces high-fidelity results via simply controlling the geometric relationship between noises without tuning or optimization. Our extensive experiments demonstrate the efficacy of our approach, highlighting its potential to advance the state-of-the-art in semantics manipulation.

文本生成扩散模型语义编辑免训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。