用低分辨率代理加速4K图像局部编辑,高效又保质。
HierEdit: Region-Aware Hierarchical Diffusion for Efficient High-Resolution Editing

- 先用低分辨率代理定位修改区域,再局部精细修复
- 4K图像编辑速度提升显著,无需专门高分辨训练数据
- 适合需要快速高质量图像编辑的设计师和开发者
高分辨率图像编辑对专业创作至关重要,但现有基于多模态扩散模型的方法计算效率低,且受限于较低分辨率。现有方法通常冗余处理整张图像或依赖大规模高分辨率数据集,导致训练与推理成本高昂。本文提出HierEdit,一种面向高效可扩展高分辨率编辑的区域感知分层扩散框架。首先利用现成编辑模型在低分辨率代理上进行编辑,生成参考并定位修改区域;随后采用分层局部窗口扩散模型(Local-Window MMDiT),仅在原图高分辨率中对修改区域进行精细化修复,同时将未改变区域作为条件输入。低分辨率代理还提供结构引导和中间去噪监督(Inference Acceleration),确保全局语义一致性和生成稳定性,无需全分辨率注意力计算。该针对性分层设计使模型能在无专用高分辨率训练数据的情况下,实现高达4K分辨率的快速、高保真编辑。大量实验表明,HierEdit在通用分辨率数据集上达到竞争力的视觉质量,显著加速推理,并可无缝扩展至超高清4K编辑。
原文摘要 · Abstract (English)
High-resolution image editing is essential for professional and creative applications, yet existing multimodal diffusion-based editors remain computationally inefficient and constrained to relatively low resolutions. Current approaches redundantly process the entire image canvas or rely on large-scale high-resolution datasets, resulting in substantial training and inference costs. We introduce HierEdit, a region-aware hierarchical diffusion framework designed for efficient and scalable high-resolution image editing. Our method first performs edits on a low-resolution proxy using an off-the-shelf editing model to generate a reference and to localize the modified regions. A hierarchical local-window diffusion model (\textbf{Local-Window MMDiT}) that refines only edited regions within the original high-res image, while reusing the unaltered regions as conditioning inputs. The low-resolution proxy further provides structural guidance and intermediate denoising supervision (\textbf{Inference Acceleration}) , ensuring consistent global semantics and stable generation without the need for full-resolution attention computation. This targeted and hierarchical design enables fast, high-fidelity editing of images up to 4K resolution without any specialized high-resolution training data. Extensive experiments demonstrate that HierEdit achieves competitive visual quality on commodity-resolution datasets while significantly accelerating inference and extending seamlessly to ultra-high-resolution 4K editing. Please check our {\href{https://peteryyzhang.github.io/HierEdit-page/}{\textbf{Project Page}}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。