用高分辨率参考图同时修复生成缺陷并超分,让AI图像更逼真可用。
RefGC-SR$^2$: Reference-guided Super-Resolution and Refinement of AI Generated Content

- 利用原始高分辨率参考图,在后处理阶段同步超分与去伪影。
- 相比现有方法,生成图像身份更忠实、细节更丰富,质量显著提升。
- 适合需要高质量图像输出的个性化生成场景,如艺术创作与设计。
参考引导生成(如物体合成、定制)发展迅速,但现有流程存在根本局限:用户提供的高分辨率参考图(HRRI)在输入模型前被下采样至固定低分辨率(LR),导致细粒度细节在输出前即丢失;随后生成过程还会引入新伪影(如身份失真)。现有参考引导内容修复(RefGCR)方法虽可修正部分伪影,但仍局限于低分辨率域;参考引导超分辨率(RefSR)方法能恢复分辨率,却假设自然图像退化,忽略生成模型带来的伪影分布。为统一解决上述问题,我们提出新任务:参考引导生成内容超分-精修(RefGC-SR²),在后处理阶段复用原始高分辨率参考图,同时恢复丢失细节、精修生成伪影并上采样。我们构建首个真实世界三元组数据生成流水线,训练双图条件生成器合成公共预训练模型无法提供的低质量锚点对。进一步提出频域感知扩散变换器模型,选择性注入来自参考图的精细细节,同时消除生成伪影。大量实验表明,所提模型成功(i)忠实还原对象身份,(ii)恢复高分辨率细节,最终结果显著优于现有RefGCR与RefSR基线,质量更高,更具实用性。
原文摘要 · Abstract (English)
Reference-guided generation (e.g., object compositing, customization) has progressed rapidly, yet current pipelines share a fundamental limitation: the object-centric high-resolution reference image (HRRI) provided by users is downsampled to a fixed low-resolution (LR) before being fed into the model, so the fine-grained details are discarded before the output is even produced. In addition, the generation step then introduces its own artifacts (e.g., identity distortion) on top of this loss. Existing reference-guided generated content refinement (RefGCR) methods can correct some of these artifacts but still operate in the LR domain; reference-guided super-resolution (RefSR) methods recover resolution but assume natural-image degradations and ignore the artifact distribution of generative pipelines. To address both gaps in a single formulation, we introduce a new task: reference-guided generated content super-resolution-refinement (RefGC-SR$^2$), where the original HRRI is reused at the post-processing stage to recover lost details, refine generative artifacts, and upscale the output simultaneously. We construct the first real-world triplet data generation pipeline for this RefGC-SR$^2$ task, training a diptych-conditioned generator to synthesize paired low-quality anchors that public pretrained models cannot provide. We further present a frequency-aware diffusion transformer model for RefGC-SR$^2$ that selectively injects fine details from the HRRI while removing generative artifacts. Extensive experiments demonstrate that our RefGC-SR$^2$ model successfully (i) refines the object identity faithfully with respect to the reference, and (ii) recovers high-resolution details, so that the final result is significantly higher quality and practically more usable compared to existing RefGCR and RefSR baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。