一歩で高速な人物画像個別化を実現する新手法
SwiftPie: Lightning-fast Subject-driven Image Personalization via One step Diffusion

- 一歩の拡散プロセスで被験者のアイデンティティを注入
- 1ステップでマルチステップ手法と同等の顔再現精度を達成
- リアルタイム対応が可能、インタラクティブな画像生成に最適
扩散模型在高质量图像生成方面取得了显著进展,激发了以图像引导的个性化生成任务的兴趣,如主体驱动的图像个性化。尽管现有方法实现了出色的个性化效果,但通常依赖于计算密集型微调、迭代优化或多步去噪过程,严重阻碍了其在实时应用中的部署与交互能力。本文提出SwiftPie,首个一步扩散图像个性化工具,可实现闪电般的个性化图像生成。SwiftPie引入了一种新型双分支身份注入机制,有效将主体身份整合进一步扩散模型中。此外,我们还采用掩码引导的缩放策略,进一步提升单步扩散中的主体上下文一致性。大量实验表明,SwiftPie不仅在个性化生成速度上表现卓越,且在身份保真度和提示对齐方面与多步方法相当。该工作为实时、高质量的个性化图像生成开辟了新路径,推动了交互式视觉合成的发展。
原文摘要 · Abstract (English)
Diffusion models have achieved remarkable success in high-quality image synthesis, sparking interest in image-guided generation tasks such as subject-driven image personalization. Despite their impressive personalization results, existing methods typically rely on computationally intensive fine-tuning, iterative optimization, or multi-step denoising processes, which significantly hinder their deployment and interactive capability in real-time applications. In this work, we present SwiftPie, the first one-step diffusion image personalization tool that enables lightning-fast generation of personalized images. SwiftPie introduces a novel dual-branch identity injection mechanism that effectively integrates subject identity into a one-step diffusion model. In addition, we incorporate a mask-guided rescaling strategy to further enhance subject contextualization within a single diffusion step. Extensive experiments demonstrate that SwiftPie not only delivers superior image personalization speed but also achieves comparable performance with multi-step approaches in both identity fidelity and prompt alignment. This work opens new opportunities for real-time, high-quality personalized image generation, paving the way for interactive visual synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。