首个统一人脸细粒度理解与生成的模型,提升细节还原与文本对齐能力。
UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
- 采用双离散扩散损失统一生成与理解任务,优化负对数似然逼近。
- 多层级专家混合架构增强属性与身份特征融合,缓解表征遗忘问题。
- 构建100万级图文问答数据集,覆盖更广人脸属性,适合高保真应用研究。
统一多模态模型(UMMs)在跨模态基础研究中展现出强大潜力,但在人脸领域仍面临两大挑战:(1)发展碎片化,现有方法未能将理解与生成统一于单一框架,阻碍通用人工智能进展;(2)缺乏细粒度面部属性,制约高保真应用。为此,我们提出首个专为细粒度人脸理解与生成设计的统一模型UniF²ace。首先,引入一种新型理论框架与双离散扩散(D3Diff)损失,将掩码生成模型与离散分数匹配扩散结合,更精确逼近负对数似然,显著提升文本引导下高质量面部细节合成能力。其次,提出多层级分组专家混合架构,自适应融合语义与身份嵌入,缓解表征演化中的属性遗忘现象。最后,构建了UniF²aceD-1M数据集,包含13万张细粒度图像-描述对和100万组视觉问答对,涵盖比现有数据集更广泛的人脸属性。大量实验表明,UniF²ace在理解与生成任务上均优于同等规模模型,分别取得7.1%更高的Desc-GPT得分和6.6%更高的VQA得分。
原文摘要 · Abstract (English)
Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and generation. However, existing research in the face domain primarily faces two challenges: $\textbf{(1)}$ $\textbf{fragmentation development}$, with existing methods failing to unify understanding and generation into a single one, hindering the way to artificial general intelligence. $\textbf{(2) lack of fine-grained facial attributes}$, which are crucial for high-fidelity applications. To handle those issues, we propose $\textbf{UniF$^2$ace}$, $\textit{the first UMM specifically tailored for fine-grained face understanding and generation}$. $\textbf{First}$, we introduce a novel theoretical framework with a Dual Discrete Diffusion (D3Diff) loss, unifying masked generative models with discrete score matching diffusion and leading to a more precise approximation of the negative log-likelihood. Moreover, this D3Diff significantly enhances the model's ability to synthesize high-fidelity facial details aligned with text input. $\textbf{Second}$, we propose a multi-level grouped Mixture-of-Experts architecture, adaptively incorporating the semantic and identity facial embeddings to complement the attribute forgotten phenomenon in representation evolvement. $\textbf{Finally}$, to this end, we construct UniF$^2$aceD-1M, a large-scale dataset comprising 130K fine-grained image-caption pairs and 1M visual question-answering pairs, spanning a much wider range of facial attributes than existing datasets. Extensive experiments demonstrate that UniF$^2$ace outperforms existing models with a similar scale in both understanding and generation tasks, with 7.1\% higher Desc-GPT and 6.6\% higher VQA-score, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。