分离身份的语义与视觉特征,实现零训练个性化生成。
DVI: Disentangling Semantic and Visual Identity for Training-Free Personalized Generation
- 将身份分解为语义与视觉两个独立流,利用VAE隐空间均值方差描述整体视觉氛围。
- 无需微调即可保持高身份保真度,且在IBench评测中优于现有方法。
- 适合需要快速个性化生成且关注环境氛围一致性的应用者。
近期无微调的身份定制方法虽能实现高面部保真度,却常忽略光照、皮肤纹理和环境色调等视觉上下文,导致语义-视觉不协调,出现‘贴纸式’的不自然效果。本文提出DVI(解耦视觉身份)框架,一种零样本方法,将身份信息正交解耦为细粒度语义流与粗粒度视觉流。不同于仅依赖语义向量的方法,DVI利用VAE隐空间的统计特性,以均值和方差作为轻量级全局视觉氛围描述符。引入无参数特征调制机制,自适应地将视觉统计信息注入语义嵌入,实现参考图像‘视觉灵魂’的无训练注入。同时设计动态时间粒度调度器,匹配扩散过程:早期优先处理视觉氛围,后期聚焦语义细节优化。大量实验表明,DVI显著提升视觉一致性与氛围保真度,无需参数微调,保持强身份保留能力,在IBench评测中超越当前最优方法。
原文摘要 · Abstract (English)
Recent tuning-free identity customization methods achieve high facial fidelity but often overlook visual context, such as lighting, skin texture, and environmental tone. This limitation leads to ``Semantic-Visual Dissonance,'' where accurate facial geometry clashes with the input's unique atmosphere, causing an unnatural ``sticker-like'' effect. We propose **DVI (Disentangled Visual-Identity)**, a zero-shot framework that orthogonally disentangles identity into fine-grained semantic and coarse-grained visual streams. Unlike methods relying solely on semantic vectors, DVI exploits the inherent statistical properties of the VAE latent space, utilizing mean and variance as lightweight descriptors for global visual atmosphere. We introduce a **Parameter-Free Feature Modulation** mechanism that adaptively modulates semantic embeddings with these visual statistics, effectively injecting the reference's ``visual soul'' without training. Furthermore, a **Dynamic Temporal Granularity Scheduler** aligns with the diffusion process, prioritizing visual atmosphere in early denoising stages while refining semantic details later. Extensive experiments demonstrate that DVI significantly enhances visual consistency and atmospheric fidelity without parameter fine-tuning, maintaining robust identity preservation and outperforming state-of-the-art methods in IBench evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。