用轻量模型+大模型引导,快速生成高保真人物图像。
Stencil: Subject-Driven Generation with Context Guidance
- 轻量模型实时微调主体图像,大模型冻结提供上下文指导。
- 1分钟内生成高质量新姿势图像,超越现有方法表现。
- 适合需要快速生成特定人物多场景图像的创作者。
当前文本到图像扩散模型在生成视觉内容方面表现出色,但难以保持主体一致性与上下文连贯性。现有微调方法存在质量与效率的权衡:微调大型模型虽提升保真度但计算成本高,微调小型模型虽高效却牺牲图像质量。此外,在少量主体图像上微调预训练模型会破坏原有先验知识,导致效果不佳。为此,我们提出Stencil框架,推理时协同使用两个扩散模型:通过轻量模型对主体图像进行高效微调,同时利用大型冻结预训练模型在推理阶段提供丰富的上下文先验,显著提升生成质量且开销极小。Stencil可在不到一分钟内生成高保真、新颖的主体图像,达到当前最佳性能,为特定主体生成设立了新基准。
原文摘要 · Abstract (English)
Recent text-to-image diffusion models can generate striking visuals from text prompts, but they often fail to maintain subject consistency across generations and contexts. One major limitation of current fine-tuning approaches is the inherent trade-off between quality and efficiency. Fine-tuning large models improves fidelity but is computationally expensive, while fine-tuning lightweight models improves efficiency but compromises image fidelity. Moreover, fine-tuning pre-trained models on a small set of images of the subject can damage the existing priors, resulting in suboptimal results. To this end, we present Stencil, a novel framework that jointly employs two diffusion models during inference. Stencil efficiently fine-tunes a lightweight model on images of the subject, while a large frozen pre-trained model provides contextual guidance during inference, injecting rich priors to enhance generation with minimal overhead. Stencil excels at generating high-fidelity, novel renditions of the subject in less than a minute, delivering state-of-the-art performance and setting a new benchmark in subject-driven generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。