arXiv:2511.19990cs.CV2025-11被引 5

用强化学习提升参考图生成的细节一致性。

OmniRefiner: Reinforcement-Guided Local Diffusion Refinement

  • 分两阶段优化:先联合输入草稿与参考图,再用强化学习增强局部编辑。
  • 在复杂修复任务上超越开源与商用模型,细节保留更精准。
  • 适合需要高保真图像修改的视觉生成场景。

参考引导的图像生成进展迅速,但现有扩散模型在使用参考图精修生成图像时仍难以保持细粒度视觉细节。这一限制源于基于VAE的潜在空间压缩会丢失细微纹理信息,导致身份与属性特异性线索消失。此外,现有后编辑方法在放大局部细节时,常导致光照、纹理或形状与原图不一致。为此,我们提出 extit{OmniRefiner},一个注重细节的精修框架,通过两次连续的参考驱动修正提升像素级一致性。首先,通过微调单图扩散编辑器,使其联合接收草稿图像与参考图像,实现全局一致的精修同时保持结构真实。随后,引入强化学习进一步强化局部编辑能力,显式优化细节准确率与语义一致性。大量实验表明, extit{OmniRefiner}显著提升参考对齐度与细粒度细节保留,在具有挑战性的参考引导修复基准测试中优于开源与商业模型。

原文摘要 · Abstract (English)

Reference-guided image generation has progressed rapidly, yet current diffusion models still struggle to preserve fine-grained visual details when refining a generated image using a reference. This limitation arises because VAE-based latent compression inherently discards subtle texture information, causing identity- and attribute-specific cues to vanish. Moreover, post-editing approaches that amplify local details based on existing methods often produce results inconsistent with the original image in terms of lighting, texture, or shape. To address this, we introduce \ourMthd{}, a detail-aware refinement framework that performs two consecutive stages of reference-driven correction to enhance pixel-level consistency. We first adapt a single-image diffusion editor by fine-tuning it to jointly ingest the draft image and the reference image, enabling globally coherent refinement while maintaining structural fidelity. We then apply reinforcement learning to further strengthen localized editing capability, explicitly optimizing for detail accuracy and semantic consistency. Extensive experiments demonstrate that \ourMthd{} significantly improves reference alignment and fine-grained detail preservation, producing faithful and visually coherent edits that surpass both open-source and commercial models on challenging reference-guided restoration benchmarks.

图像精修扩散模型强化学习细节保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。