arXiv:2510.10156cs.CV2025-10

统一人物生成与编辑,实现身份一致且可控的图像创作。

ReMix: Towards a Unified View of Consistent Character Generation and Editing

  • 用多模态大模型编辑语义特征,无需微调即可适配扩散模型。
  • 通过共享噪声空间联合去噪,保持姿态与像素级一致性。
  • 适合个性化生成、风格迁移等需要身份稳定的场景。

近期大规模文本到图像扩散模型(如FLUX.1)显著提升了人物生成与编辑的视觉质量,但现有方法极少在单一框架中统一这两项任务。基于生成的方法在不同实例间难以保持细粒度身份一致,而基于编辑的方法常丧失空间控制力和指令对齐性。为此,我们提出ReMix,一个统一的人物一致性生成与编辑框架。其核心包含两个组件:ReMix模块与IP-ControlNet。ReMix模块利用多模态大模型(MLLM)的推理能力,编辑输入图像的语义特征,并将指令嵌入适配至原生DiT骨干网络,无需微调。这确保了语义布局的一致性,但像素级一致性和姿态可控性仍具挑战。为此,IP-ControlNet扩展ControlNet,将参考图像的语义与布局线索解耦,并引入ε-等变潜在空间,在共享噪声空间中联合去噪参考图与目标图。受收敛演化与量子退相干启发——环境噪声促使状态趋同——该设计促进隐空间中的特征对齐,实现一致对象生成的同时保留身份特征。ReMix支持个性化生成、图像编辑、风格迁移及多条件合成等多种任务。大量实验验证了其作为统一框架在人物一致性图像生成与编辑中的有效性与高效性。

原文摘要 · Abstract (English)

Recent advances in large-scale text-to-image diffusion models (e.g., FLUX.1) have greatly improved visual fidelity in consistent character generation and editing. However, existing methods rarely unify these tasks within a single framework. Generation-based approaches struggle with fine-grained identity consistency across instances, while editing-based methods often lose spatial controllability and instruction alignment. To bridge this gap, we propose ReMix, a unified framework for character-consistent generation and editing. It constitutes two core components: the ReMix Module and IP-ControlNet. The ReMix Module leverages the multimodal reasoning ability of MLLMs to edit semantic features of input images and adapt instruction embeddings to the native DiT backbone without fine-tuning. While this ensures coherent semantic layouts, pixel-level consistency and pose controllability remain challenging. To address this, IP-ControlNet extends ControlNet to decouple semantic and layout cues from reference images and introduces an ε-equivariant latent space that jointly denoises the reference and target images within a shared noise space. Inspired by convergent evolution and quantum decoherence,i.e., where environmental noise drives state convergence, this design promotes feature alignment in the hidden space, enabling consistent object generation while preserving identity. ReMix supports a wide range of tasks, including personalized generation, image editing, style transfer, and multi-condition synthesis. Extensive experiments validate its effectiveness and efficiency as a unified framework for character-consistent image generation and editing.

图像生成身份一致扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。