不微调也能精准保留人物身份,避免背景干扰。
Inject Where It Matters: Training-Free Spatially-Adaptive Identity Preservation for Text-to-Image Personalization

- 用注意力响应生成空间掩码,区分人脸与非人脸区域
- 在IBench上文本一致性和身份保真度分别达0.281和0.827
- 适合需要快速个性化生成的图像创作场景
个性化文生图旨在将特定身份融入任意场景。现有免微调方法多采用空间均匀注入,导致身份特征污染非面部区域(如背景和光照),降低文本契合度。为解决此问题且无需昂贵微调,本文提出SpatialID——一种训练自由的空间自适应身份调制框架。该方法基于交叉注意力响应构建空间掩码提取器,从根本上将身份注入分解为仅与人脸相关和与上下文无关的区域。同时引入时序-空间调度策略,动态调整空间约束:从高斯先验逐步过渡到注意力掩码,并实现自适应松弛,以匹配扩散生成过程。在IBench上的大量实验表明,SpatialID在文本一致性(CLIP-T: 0.281)、视觉一致性(CLIP-I: 0.827)和图像质量(IQ: 0.523)方面均达到当前最优,显著消除背景污染,同时保持强身份保留能力。
原文摘要 · Abstract (English)
Personalized text-to-image generation aims to integrate specific identities into arbitrary contexts. However, existing tuning-free methods typically employ Spatially Uniform Visual Injection, causing identity features to contaminate non-facial regions (e.g., backgrounds and lighting) and degrading text adherence. To address this without expensive fine-tuning, we propose SpatialID, a training-free spatially-adaptive identity modulation framework. SpatialID fundamentally decouples identity injection into face-relevant and context-free regions using a Spatial Mask Extractor derived from cross-attention responses. Furthermore, we introduce a Temporal-Spatial Scheduling strategy that dynamically adjusts spatial constraints - transitioning from Gaussian priors to attention-based masks and adaptive relaxation - to align with the diffusion generation dynamics. Extensive experiments on IBench demonstrate that SpatialID achieves state-of-the-art performance in text adherence (CLIP-T: 0.281), visual consistency (CLIP-I: 0.827), and image quality (IQ: 0.523), significantly eliminating background contamination while maintaining robust identity preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。