通过频率感知的去噪机制,实现文本引导图像编辑的精准局部修改。
FDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image Editing
- 利用小波分解图像,在不同频段选择性优化局部区域。
- 在多个数据集上显著减少细节丢失与颜色偏差,提升编辑精度。
- 适用于2D图像与3D纹理编辑,适合需要精细控制的生成任务。
使用文本到图像(T2I)模型进行文本引导的图像编辑常导致不理想结果,如局部细节丢失和颜色变化。本文分析失败原因,发现源于对所有频率带的无差别优化,而实际上仅特定频率需调整。为此,提出一种简单有效的方法,可在局部空间区域中选择性优化特定频率带,实现精准编辑。方法基于小波变换将图像分解为多频段、多尺度表示,支持不同层次细节的精确修改。进一步对比了多种频域技术并扩展至3D纹理编辑:在三平面表示上进行频率分解,实现3D纹理的频率感知调整。定量评估与用户研究均证明该方法能生成高质量且精确的编辑结果。
原文摘要 · Abstract (English)
Text-guided image editing using Text-to-Image (T2I) models often fails to yield satisfactory results, frequently introducing unintended modifications, such as the loss of local detail and color changes. In this paper, we analyze these failure cases and attribute them to the indiscriminate optimization across all frequency bands, even though only specific frequencies may require adjustment. To address this, we introduce a simple yet effective approach that enables the selective optimization of specific frequency bands within localized spatial regions for precise edits. Our method leverages wavelets to decompose images into different spatial resolutions across multiple frequency bands, enabling precise modifications at various levels of detail. To extend the applicability of our approach, we provide a comparative analysis of different frequency-domain techniques. Additionally, we extend our method to 3D texture editing by performing frequency decomposition on the triplane representation, enabling frequency-aware adjustments for 3D textures. Quantitative evaluations and user studies demonstrate the effectiveness of our method in producing high-quality and precise edits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。