arXiv:2602.03448cs.CVcs.AI2026-02被引 1

通过分层引导实现多主体图像生成的身份一致性与精准控制。

Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation

  • 分层设计:概念层用VAE丢弃增强语义依赖,外观层用注意力约束匹配区域
  • 在多个数据集上显著提升身份一致性和文本指令遵循能力
  • 适合需要高精度多角色图像合成的研究者和开发者

多主体图像生成旨在合成忠实保留多个参考主体身份并遵循文本指令的图像。现有方法常因扩散模型隐式关联文本与参考图像而出现身份不一致和组合控制有限的问题。本文提出分层概念到外观引导(CAG)框架,从高层概念到细粒度外观提供显式结构化监督。概念层引入VAE丢弃训练策略,随机剔除参考VAE特征,促使模型更依赖视觉语言模型(VLM)的鲁棒语义信号,从而在缺乏完整外观线索时仍能保持概念层面的一致性生成。外观层将VLM推导的对应关系融入扩散Transformer(DiT)中的对应感知掩码注意力模块,使每个文本标记仅关注其匹配的参考区域,确保属性绑定精确且多主体组合可靠。大量实验表明,本方法在多主体图像生成任务中达到当前最优性能,显著提升提示遵循度与主体一致性。

原文摘要 · Abstract (English)

Multi-subject image generation aims to synthesize images that faithfully preserve the identities of multiple reference subjects while following textual instructions. However, existing methods often suffer from identity inconsistency and limited compositional control, as they rely on diffusion models to implicitly associate text prompts with reference images. In this work, we propose Hierarchical Concept-to-Appearance Guidance (CAG), a framework that provides explicit, structured supervision from high-level concepts to fine-grained appearances. At the conceptual level, we introduce a VAE dropout training strategy that randomly omits reference VAE features, encouraging the model to rely more on robust semantic signals from a Visual Language Model (VLM) and thereby promoting consistent concept-level generation in the absence of complete appearance cues. At the appearance level, we integrate the VLM-derived correspondences into a correspondence-aware masked attention module within the Diffusion Transformer (DiT). This module restricts each text token to attend only to its matched reference regions, ensuring precise attribute binding and reliable multi-subject composition. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the multi-subject image generation, substantially improving prompt following and subject consistency.

多主体生成扩散模型视觉语言模型身份一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。