arXiv:2503.15406cs.CV2025-03CVPR被引 9

用一张生活照生成多种场景下的人物图像,保留全身特征。

Visual Persona: Foundation Model for Full-Body Human Customization

  • 基于视觉语言模型构建数据筛选流程,获取高质量配对图像
  • 训练出58万张跨10万身份的全身体型数据集,支持多样化生成
  • 通过区域编码与嵌入投影,精准迁移全身外观特征

我们提出Visual Persona,一个用于文本到图像全身体型定制的基础模型。给定一张自然环境中的个人图像,该模型可依据文本描述生成多样化的该人物图像,涵盖身体结构与场景变化。不同于以往仅关注面部识别的方法,本方法捕捉详细的全身外观特征并与其文本描述一致。训练该模型需大规模成对的人体数据,包含每位个体的多张保持一致全身体型的图像,但此类数据极难获取。为此,我们提出一种利用视觉-语言模型评估全身外观一致性的数据整理流程,构建了包含58万张配对图像、覆盖10万唯一身份的Visual Persona-500K数据集。为实现精确外观迁移,我们设计了一种适配预训练文本到图像扩散模型的变压器编码器-解码器架构,将输入图像划分为不同身体区域,分别编码为局部外观特征,并独立投影为密集身份嵌入,以条件化扩散模型生成定制化图像。Visual Persona在多个基准上持续优于现有方法,能从自然图像中生成高质量定制图像。大量消融实验验证了设计选择的有效性,并展示了其在多种下游任务中的通用性。

原文摘要 · Abstract (English)

We introduce Visual Persona, a foundation model for text-to-image full-body human customization that, given a single in-the-wild human image, generates diverse images of the individual guided by text descriptions. Unlike prior methods that focus solely on preserving facial identity, our approach captures detailed full-body appearance, aligning with text descriptions for body structure and scene variations. Training this model requires large-scale paired human data, consisting of multiple images per individual with consistent full-body identities, which is notoriously difficult to obtain. To address this, we propose a data curation pipeline leveraging vision-language models to evaluate full-body appearance consistency, resulting in Visual Persona-500K, a dataset of 580k paired human images across 100k unique identities. For precise appearance transfer, we introduce a transformer encoder-decoder architecture adapted to a pre-trained text-to-image diffusion model, which augments the input image into distinct body regions, encodes these regions as local appearance features, and projects them into dense identity embeddings independently to condition the diffusion model for synthesizing customized images. Visual Persona consistently surpasses existing approaches, generating high-quality, customized images from in-the-wild inputs. Extensive ablation studies validate design choices, and we demonstrate the versatility of Visual Persona across various downstream tasks.

人物定制扩散模型视觉人格

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。