用单图快速生成个性化图像,效率比现有方法提升显著
HybridBooth: Hybrid Prompt Inversion for Efficient Subject-Driven Generation
- 结合优化与直接回归,分两阶段生成视觉概念的文本嵌入
- 仅需一张图像即可完成高效反演,且保持模型泛化能力
- 适合需要快速个性化生成的设计师、创意工作者
文本到图像扩散模型在文本提示下展现出强大的创造力,但基于特定主体的个性化生成(即主题驱动生成)仍具挑战。为此,我们提出一种名为HybridBooth的新混合框架,融合基于优化和直接回归方法的优势。该框架分为两个阶段:首先通过微调编码器进行词嵌入探测,生成稳健的初始词嵌入;其次通过优化关键参数,进一步适配特定主体图像,实现编码器的精细化调整。该方法可在仅一张图像的前提下,高效且准确地将视觉概念反演为文本嵌入,同时保持模型的通用性。
原文摘要 · Abstract (English)
Recent advancements in text-to-image diffusion models have shown remarkable creative capabilities with textual prompts, but generating personalized instances based on specific subjects, known as subject-driven generation, remains challenging. To tackle this issue, we present a new hybrid framework called HybridBooth, which merges the benefits of optimization-based and direct-regression methods. HybridBooth operates in two stages: the Word Embedding Probe, which generates a robust initial word embedding using a fine-tuned encoder, and the Word Embedding Refinement, which further adapts the encoder to specific subject images by optimizing key parameters. This approach allows for effective and fast inversion of visual concepts into textual embedding, even from a single image, while maintaining the model's generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。