通过重新生成扩大修改空间,显著提升统一多模态模型的图像精修效果。
Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models

- 将精修任务转为条件重生成,不再依赖编辑指令和严格保留像素
- 在多个基准上实现显著提升,最高达77.41分(原61.53)
- 适合需要高质量图像生成与精准语义对齐的研究者
统一多模态模型(UMMs)将视觉理解与生成整合于单一框架中。对于文本到图像(T2I)任务,这种统一能力使UMMs可在初始生成后进行精修,有望突破性能上限。现有基于UMM的精修方法主要采用精修-编辑(RvE)范式,即生成编辑指令以修正不一致区域,同时保留一致内容。然而,编辑指令往往仅粗略描述提示与图像之间的错位,导致精修不完整;且像素级保留虽必要于编辑,却无谓限制了精修的有效修改空间。为此,我们提出精修-重生成(RvR)框架,将精修重构为条件图像重生成,而非编辑。RvR不依赖编辑指令,也不强制内容保留,而是以目标提示和初始图像的语义标记为条件进行图像重生成,从而实现更完整的语义对齐并扩大修改空间。大量实验表明,RvR显著提升性能:Geneval从0.78提升至0.91,DPGBench从84.02提升至87.21,UniGenBench++从61.53提升至77.41。
原文摘要 · Abstract (English)
Unified multimodal models (UMMs) integrate visual understanding and generation within a single framework. For text-to-image (T2I) tasks, this unified capability allows UMMs to refine outputs after their initial generation, potentially extending the performance upper bound. Current UMM-based refinement methods primarily follow a refinement-via-editing (RvE) paradigm, where UMMs produce editing instructions to modify misaligned regions while preserving aligned content. However, editing instructions often describe prompt-image misalignment only coarsely, leading to incomplete refinement. Moreover, pixel-level preservation, though necessary for editing, unnecessarily restricts the effective modification space for refinement. To address these limitations, we propose Refinement via Regeneration (RvR), a novel framework that reformulates refinement as conditional image regeneration rather than editing. Instead of relying on editing instructions and enforcing strict content preservation, RvR regenerates images conditioned on the target prompt and the semantic tokens of the initial image, enabling more complete semantic alignment with a larger modification space. Extensive experiments demonstrate the effectiveness of RvR, improving Geneval from 0.78 to 0.91, DPGBench from 84.02 to 87.21, and UniGenBench++ from 61.53 to 77.41.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。