无需训练,用参考图直接生成主题一致的图像。
IP-Prompter: Training-Free Theme-Specific Image Generation via Dynamic Visual Prompting
- 用参考图作为视觉提示,动态优化生成内容。
- 在角色一致性、风格统一性上优于现有方法。
- 适合角色设计、故事连贯生成等场景。
成长中吸引我们的故事与角色塑造了独特的幻想世界,图像成为体验这些世界的主媒介。当前文本到图像生成中,通过主题特定数据微调模型已成为主流,但主题生成涉及角色、场景、物体等多样元素,带来如何自适应生成多角色、多概念、连续主题图像(TSI)的挑战。且微调成本高、易过拟合。本文提出核心问题:能否让生成模型像大语言模型使用文本上下文一样,直接利用图像作为上下文?为此,我们提出IP-Prompter——一种无需训练的主题图像生成方法。其引入视觉提示机制,将参考图像融入生成模型,用户无需额外训练即可指定目标主题。为进一步提升效果,提出动态视觉提示(DVP)机制,通过迭代优化视觉提示,提高生成图像的准确性和质量。该方法支持连贯故事生成、角色设计、真实角色生成及风格引导图像生成。对比实验表明,IP-Prompter在角色身份保持、风格一致性、文本对齐方面显著优于现有个性化方法,为主题图像生成提供强大灵活的解决方案。
原文摘要 · Abstract (English)
The stories and characters that captivate us as we grow up shape unique fantasy worlds, with images serving as the primary medium for visually experiencing these realms. Personalizing generative models through fine-tuning with theme-specific data has become a prevalent approach in text-to-image generation. However, unlike object customization, which focuses on learning specific objects, theme-specific generation encompasses diverse elements such as characters, scenes, and objects. Such diversity also introduces a key challenge: how to adaptively generate multi-character, multi-concept, and continuous theme-specific images (TSI). Moreover, fine-tuning approaches often come with significant computational overhead, time costs, and risks of overfitting. This paper explores a fundamental question: Can image generation models directly leverage images as contextual input, similarly to how large language models use text as context? To address this, we present IP-Prompter, a novel training-free TSI generation method. IP-Prompter introduces visual prompting, a mechanism that integrates reference images into generative models, allowing users to seamlessly specify the target theme without requiring additional training. To further enhance this process, we propose a Dynamic Visual Prompting (DVP) mechanism, which iteratively optimizes visual prompts to improve the accuracy and quality of generated images. Our approach enables diverse applications, including consistent story generation, character design, realistic character generation, and style-guided image generation. Comparative evaluations against state-of-the-art personalization methods demonstrate that IP-Prompter achieves significantly better results and excels in maintaining character identity preserving, style consistency and text alignment, offering a robust and flexible solution for theme-specific image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。