4K级图像修复新框架,兼顾全局结构与局部细节。
Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency Adapter
- 分两阶段适配:先建全局结构,再精细修复局部。
- 在OpenImages和Photo-Concept-Bucket上超越现有方法。
- 适合需要高分辨率、高保真修复的视觉生成任务。
本文提出Patch-Adapter,一种高效实现高分辨率文本引导图像修复的框架。相比现有方法仅限于低分辨率,该方法可达到4K以上分辨率,同时保持精确的内容一致性和提示对齐,这两项挑战在分辨率提升和纹理复杂度增加时尤为显著。Patch-Adapter采用双阶段适配架构,将扩散模型分辨率从1K扩展至4K+,无需结构性重构:(1) 双上下文适配器在低分辨率下学习掩码与未掩码区域的一致性,建立全局结构一致性;(2) 参考块适配器引入块级注意力机制,在全分辨率下进行修复,通过自适应特征融合保持局部细节保真度。该双阶段架构通过解耦全局语义与局部细化,独特地解决了高分辨率修复中的可扩展性问题。实验表明,Patch-Adapter不仅有效消除大规模修复中的伪影,还在OpenImages和Photo-Concept-Bucket数据集上达到最先进性能,显著优于现有方法在感知质量和文本提示遵循度上的表现。
原文摘要 · Abstract (English)
In this work, we present Patch-Adapter, an effective framework for high-resolution text-guided image inpainting. Unlike existing methods limited to lower resolutions, our approach achieves 4K+ resolution while maintaining precise content consistency and prompt alignment, two critical challenges in image inpainting that intensify with increasing resolution and texture complexity. Patch-Adapter leverages a two-stage adapter architecture to scale the diffusion model's resolution from 1K to 4K+ without requiring structural overhauls: (1) Dual Context Adapter learns coherence between masked and unmasked regions at reduced resolutions to establish global structural consistency; and (2) Reference Patch Adapter implements a patch-level attention mechanism for full-resolution inpainting, preserving local detail fidelity through adaptive feature fusion. This dual-stage architecture uniquely addresses the scalability gap in high-resolution inpainting by decoupling global semantics from localized refinement. Experiments demonstrate that Patch-Adapter not only resolves artifacts common in large-scale inpainting but also achieves state-of-the-art performance on the OpenImages and Photo-Concept-Bucket datasets, outperforming existing methods in both perceptual quality and text-prompt adherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。