arXiv:2601.05572cs.CV2026-01

让统一多模态模型能准确处理多图编辑,保持视觉一致性和跨图信息清晰。

Towards Generalized Multi-Image Editing for Unified Multimodal Models

  • 用可学习的潜在分离器区分多图输入,实现精准条件控制。
  • 采用正弦索引编码,支持任意数量图像输入且效果稳定。
  • 构建高质量基准数据集,验证了方法在一致性与泛化上的优势。

统一多模态模型(UMMs)融合多模态理解与生成能力,但在参考多张输入图像时,难以保持视觉一致性并清晰解析视觉线索。本文提出一种可扩展的多图编辑框架,显式区分图像身份,并支持可变输入数量。算法上引入两项创新:1)可学习的潜在分离器在潜在空间中明确区分每张参考图像,实现精确且解耦的条件建模;2)正弦索引编码为同源图像的视觉标记分配连续的正弦索引嵌入,既保留图像身份,又支持对不同数量输入的泛化与外推。为促进训练与评估,我们采用逆向数据集构建方法,确保输出无伪影且可实现。实验表明,在多种多图编辑任务中,本方法在语义一致性、视觉保真度及跨图整合方面显著优于基线,验证了其在一致性和泛化能力上的优势。

原文摘要 · Abstract (English)

Unified Multimodal Models (UMMs) integrate multimodal understanding and generation, yet they are limited to maintaining visual consistency and disambiguating visual cues when referencing details across multiple input images. In this work, we propose a scalable multi-image editing framework for UMMs that explicitly distinguishes image identities and generalizes to variable input counts. Algorithmically, we introduce two innovations: 1) The learnable latent separators explicitly differentiate each reference image in the latent space, enabling accurate and disentangled conditioning. 2) The sinusoidal index encoding assigns visual tokens from the same image a continuous sinusoidal index embedding, which provides explicit image identity while allowing generalization and extrapolation on a variable number of inputs. To facilitate training and evaluation, we establish a high-fidelity benchmark using an inverse dataset construction methodology to guarantee artifact-free, achievable outputs. Experiments show clear improvements in semantic consistency, visual fidelity, and cross-image integration over prior baselines on diverse multi-image editing tasks, validating our advantages on consistency and generalization ability.

多图编辑统一模型视觉一致性泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。