Omni-ID用生成目标建模人脸整体特征,提升跨姿态表情的生成质量。
Omni-ID: Holistic Identity Representation Designed for Generative Tasks
- 通过少图多重建训练,将无结构图像转为有序身份表征
- 在多视角人脸数据集上显著优于CLIP、ArcFace等传统表示
- 适合需要高保真人脸生成的应用,如虚拟形象、视频合成
我们提出Omni-ID,一种专为生成任务设计的人脸表征。它在固定尺寸下编码个体在不同表情和姿态下的整体外观信息,将数量不一的非结构化输入图像整合为包含全局或局部身份特征的结构化表示。采用少量输入图像重建多个目标图像的训练范式,结合多解码器框架,利用不同解码器的互补优势。与通常通过判别或对比目标学习的CLIP、ArcFace等传统表征不同,Omni-ID以生成目标优化,实现更全面、细腻的身份捕捉。在自建的MFHQ多视角人脸图像数据集上训练后,其在各类生成任务中表现显著优于传统表征。
原文摘要 · Abstract (English)
We introduce Omni-ID, a novel facial representation designed specifically for generative tasks. Omni-ID encodes holistic information about an individual's appearance across diverse expressions and poses within a fixed-size representation. It consolidates information from a varied number of unstructured input images into a structured representation, where each entry represents certain global or local identity features. Our approach uses a few-to-many identity reconstruction training paradigm, where a limited set of input images is used to reconstruct multiple target images of the same individual in various poses and expressions. A multi-decoder framework is further employed to leverage the complementary strengths of diverse decoders during training. Unlike conventional representations, such as CLIP and ArcFace, which are typically learned through discriminative or contrastive objectives, Omni-ID is optimized with a generative objective, resulting in a more comprehensive and nuanced identity capture for generative tasks. Trained on our MFHQ dataset -- a multi-view facial image collection, Omni-ID demonstrates substantial improvements over conventional representations across various generative tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。