XVerse实现多主体图像生成中身份与属性的精准独立控制
XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation
- 用参考图生成令牌级文本流偏移量,实现主体特征独立调节
- 支持多主体高保真生成,避免属性混淆与图像失真
- 适合个性化图像生成与复杂场景编辑,尤其擅长多角色控制
在文本到图像生成中,对多个主体的身份和语义属性(姿态、风格、光照)进行细粒度控制时,常导致扩散变换器(DiT)的可编辑性和图像一致性下降。现有方法易引入伪影或出现属性纠缠。为此,我们提出新型多主体可控生成模型XVerse。通过将参考图像转换为针对特定标记的文本流偏移量,实现对特定主体的精确且独立控制,同时不干扰图像潜在表示或特征。因此,XVerse可在保持图像质量的前提下,提供高保真、可编辑的多主体图像合成,并具备对个体主体特征和语义属性的强控能力。该技术显著提升了个性化及复杂场景生成的能力。
原文摘要 · Abstract (English)
Achieving fine-grained control over subject identity and semantic attributes (pose, style, lighting) in text-to-image generation, particularly for multiple subjects, often undermines the editability and coherence of Diffusion Transformers (DiTs). Many approaches introduce artifacts or suffer from attribute entanglement. To overcome these challenges, we propose a novel multi-subject controlled generation model XVerse. By transforming reference images into offsets for token-specific text-stream modulation, XVerse allows for precise and independent control for specific subject without disrupting image latents or features. Consequently, XVerse offers high-fidelity, editable multi-subject image synthesis with robust control over individual subject characteristics and semantic attributes. This advancement significantly improves personalized and complex scene generation capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。