arXiv:2602.14186cs.CV2026-02被引 5

统一单图与多图编辑,提升多参考图像生成的一致性。

UniRef-Image-Edit: Towards Scalable and Consistent Multi-Reference Image Editing

  • 用序列化潜空间融合机制,动态整合多参考图
  • 分阶段训练使图像质量从1024²逐步提升至2048²
  • 首个专用于多参考生成的强化学习框架,增强一致性

我们提出UniRef-Image-Edit,一个高性能多模态生成系统,将单图编辑与多图合成统一于同一框架。现有基于扩散模型的编辑方法常因参考输入间交互有限而难以保持一致性。为此,我们引入序列扩展潜空间融合(SELF),将多张参考图动态序列化为连贯的潜空间序列。在专门的训练阶段,所有参考图在固定长度序列下受全局像素预算约束联合优化。在此基础上,我们设计两阶段训练框架:监督微调(SFT)阶段联合训练单图编辑与多图合成任务,建立稳健生成先验;采用渐进式序列长度训练策略,初始输入总像素预算为1024²,逐步增至1536²和2048²,以提升视觉保真度与跨参考一致性。该渐进式解压机制使模型逐步捕捉更精细细节,同时保持参考间稳定对齐。在强化学习(RL)阶段,我们提出多源GRPO(MSGRPO),据我们所知是首个专为多参考图像生成设计的强化学习框架,通过优化冲突视觉约束显著提升组合一致性。代码、模型、训练数据及奖励数据将开源,供社区研究使用。

原文摘要 · Abstract (English)

We present UniRef-Image-Edit, a high-performance multi-modal generation system that unifies single-image editing and multi-image composition within a single framework. Existing diffusion-based editing methods often struggle to maintain consistency across multiple conditions due to limited interaction between reference inputs. To address this, we introduce Sequence-Extended Latent Fusion (SELF), a unified input representation that dynamically serializes multiple reference images into a coherent latent sequence. During a dedicated training stage, all reference images are jointly constrained to fit within a fixed-length sequence under a global pixel-budget constraint. Building upon SELF, we propose a two-stage training framework comprising supervised fine-tuning (SFT) and reinforcement learning (RL). In the SFT stage, we jointly train on single-image editing and multi-image composition tasks to establish a robust generative prior. We adopt a progressive sequence length training strategy, in which all input images are initially resized to a total pixel budget of $1024^2$, and are then gradually increased to $1536^2$ and $2048^2$ to improve visual fidelity and cross-reference consistency. This gradual relaxation of compression enables the model to incrementally capture finer visual details while maintaining stable alignment across references. For the RL stage, we introduce Multi-Source GRPO (MSGRPO), to our knowledge the first reinforcement learning framework tailored for multi-reference image generation. MSGRPO optimizes the model to reconcile conflicting visual constraints, significantly enhancing compositional consistency. We will open-source the code, models, training data, and reward data for community research purposes.

图像编辑多参考生成扩散模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。