arXiv:2503.12590cs.CV2025-03被引 23

用简单替换法让扩散Transformer零样本实现个性化图像生成

Personalize Anything for Free with Diffusion Transformer

  • 用参考主体的特征令牌替换扩散过程中的去噪令牌
  • 在保持身份一致的同时提升图像结构多样性
  • 无需训练即可支持布局控制、多主体和掩码编辑

个性化图像生成旨在生成用户指定概念的图像并支持灵活编辑。现有无训练方法虽计算效率高,但在身份保留、适用性和与扩散Transformer(DiT)兼容性方面存在不足。本文发现,仅将去噪令牌替换为参考主体的令牌,即可实现零样本主体重建,揭示了DiT的未开发潜力。基于此,我们提出 extbf{Personalize Anything}:通过时间步自适应令牌替换,在早期注入增强身份一致性,晚期正则化提升灵活性;结合补丁扰动策略提升结构多样性。该方法无缝支持布局引导生成、多主体个性化及掩码控制编辑。评估表明其在身份保留和多功能性上达到领先水平。本工作为DiT提供了新见解,并构建了高效的个性化实用范式。

原文摘要 · Abstract (English)

Personalized image generation aims to produce images of user-specified concepts while enabling flexible editing. Recent training-free approaches, while exhibit higher computational efficiency than training-based methods, struggle with identity preservation, applicability, and compatibility with diffusion transformers (DiTs). In this paper, we uncover the untapped potential of DiT, where simply replacing denoising tokens with those of a reference subject achieves zero-shot subject reconstruction. This simple yet effective feature injection technique unlocks diverse scenarios, from personalization to image editing. Building upon this observation, we propose \textbf{Personalize Anything}, a training-free framework that achieves personalized image generation in DiT through: 1) timestep-adaptive token replacement that enforces subject consistency via early-stage injection and enhances flexibility through late-stage regularization, and 2) patch perturbation strategies to boost structural diversity. Our method seamlessly supports layout-guided generation, multi-subject personalization, and mask-controlled editing. Evaluations demonstrate state-of-the-art performance in identity preservation and versatility. Our work establishes new insights into DiTs while delivering a practical paradigm for efficient personalization.

图像生成扩散模型个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。