让图像生成保持主体一致,不用调参也不用参考图。
CoCoIns: Consistent Subject Generation via Contrastive Instantiated Concepts
- 用伪词映射潜变量,实现多次生成时主体一致
- 对比学习让模型区分不同提示与潜变量组合
- 适合需要连续图像输出的场景,如故事画册
尽管文本到图像生成模型能合成多样且逼真的内容,但多次生成中主体变化限制了其在长篇内容生成中的应用。现有方法需耗时微调、为每个主体提供参考图或依赖先前生成内容。我们提出对比概念实例化(CoCoIns),一种可在多次独立生成中保持主体一致性的框架。该框架包含生成模型和映射网络,将输入潜变量转换为特定概念实例的伪词。用户通过重复使用相同潜变量即可生成一致主体。为构建此类关联,我们设计了一种对比学习方法,训练网络区分不同提示与潜变量组合。在单主体人脸生成上的大量评估表明,CoCoIns性能接近现有方法,同时更具灵活性。我们还展示了其扩展至多主体及其他物体类别的潜力。
原文摘要 · Abstract (English)
While text-to-image generative models can synthesize diverse and faithful content, subject variation across multiple generations limits their application to long-form content generation. Existing approaches require time-consuming fine-tuning, reference images for all subjects, or access to previously generated content. We introduce Contrastive Concept Instantiation (CoCoIns), a framework that effectively synthesizes consistent subjects across multiple independent generations. The framework consists of a generative model and a mapping network that transforms input latent codes into pseudo-words associated with specific concept instances. Users can generate consistent subjects by reusing the same latent codes. To construct such associations, we propose a contrastive learning approach that trains the network to distinguish between different combinations of prompts and latent codes. Extensive evaluations on human faces with a single subject show that CoCoIns performs comparably to existing methods while maintaining greater flexibility. We also demonstrate the potential for extending CoCoIns to multiple subjects and other object categories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。