arXiv:2603.23010cs.CV2026-03

无需训练即可快速个性化任意物体,一次推理完成

Zero-Shot Personalization of Objects via Textual Inversion

  • 用学习网络预测物体专属文本反演嵌入
  • 单次前向传播实现零样本个性化生成
  • 适用于任意物体,适合快速定制场景

近期文本到图像扩散模型的进展显著提升了图像定制质量,实现了高度逼真的图像合成。然而,实现实用场景中的快速高效个性化仍是关键挑战。现有方法主要通过将身份特定嵌入注入扩散模型来加速人物个性化,但对任意物体类别泛化能力差,适用性受限。为此,我们提出一种新框架,利用学习网络预测物体特定的文本反演嵌入,并将其集成到扩散模型的UNet时间步中,实现文本条件下的个性化定制。该设计可实现单一前向传播下的快速、零样本个性化,兼具灵活性与可扩展性。大量实验在多个任务和设置中验证了方法的有效性,凸显其在支持快速、多样且包容的图像定制方面的潜力。据我们所知,这是首次在扩散模型中实现通用、无需训练的个性化,为未来个性化图像生成研究铺平道路。

原文摘要 · Abstract (English)

Recent advances in text-to-image diffusion models have substantially improved the quality of image customization, enabling the synthesis of highly realistic images. Despite this progress, achieving fast and efficient personalization remains a key challenge, particularly for real-world applications. Existing approaches primarily accelerate customization for human subjects by injecting identity-specific embeddings into diffusion models, but these strategies do not generalize well to arbitrary object categories, limiting their applicability. To address this limitation, we propose a novel framework that employs a learned network to predict object-specific textual inversion embeddings, which are subsequently integrated into the UNet timesteps of a diffusion model for text-conditional customization. This design enables rapid, zero-shot personalization of a wide range of objects in a single forward pass, offering both flexibility and scalability. Extensive experiments across multiple tasks and settings demonstrate the effectiveness of our approach, highlighting its potential to support fast, versatile, and inclusive image customization. To the best of our knowledge, this work represents the first attempt to achieve such general-purpose, training-free personalization within diffusion models, paving the way for future research in personalized image generation.

文本生成个性化扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。