不用训练,用拼图法让Flux模型生成带特定主体的图像
Flux Already Knows -- Activating Subject-Driven Image Generation without Training
- 把主体图复制成拼图布局,激活模型的身份保留能力
- 在多个评测中优于基线,人类偏好测试也表现更优
- 适合需要快速定制图像的轻量级应用
我们提出一种无需训练的零样本框架,利用基础的Flux模型实现主体驱动的图像生成。通过将任务定义为基于网格的图像补全,并简单地将主体图像以马赛克方式重复排列,即可激活强大的身份保留能力,无需额外数据、训练或推理时微调。这一“免费午餐”方法进一步通过新型级联注意力设计和元提示技术增强,提升了生成图像的保真度与多样性。实验结果表明,该方法在多个关键指标和人工偏好测试中均优于基线,尽管在某些方面存在权衡。此外,它支持多种编辑操作,包括标志插入、虚拟试穿以及主体替换或插入。结果表明,预训练的文本到图像模型可实现高质量、低资源消耗的主体驱动生成,为下游应用中的轻量级定制开辟新可能。
原文摘要 · Abstract (English)
We propose a simple yet effective zero-shot framework for subject-driven image generation using a vanilla Flux model. By framing the task as grid-based image completion and simply replicating the subject image(s) in a mosaic layout, we activate strong identity-preserving capabilities without any additional data, training, or inference-time fine-tuning. This "free lunch" approach is further strengthened by a novel cascade attention design and meta prompting technique, boosting fidelity and versatility. Experimental results show that our method outperforms baselines across multiple key metrics in benchmarks and human preference studies, with trade-offs in certain aspects. Additionally, it supports diverse edits, including logo insertion, virtual try-on, and subject replacement or insertion. These results demonstrate that a pre-trained foundational text-to-image model can enable high-quality, resource-efficient subject-driven generation, opening new possibilities for lightweight customization in downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。