arXiv:2510.18345cs.CV2025-10

用网络数据训练人脸语言模型,实现可控生成与多任务应用

GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data

  • 通过网页抓取人脸图文数据,自监督预训练跨模态模型
  • 在属性分类和表情识别上达到顶尖模型水平
  • 支持人脸编辑、换脸、去口罩等可控生成任务

相较于自然图像理解领域预训练模型的蓬勃发展,面向人脸知识学习的大规模预训练模型研究仍较有限。现有方法主要依赖人工构建和标注的人脸数据集,但标注成本高,模型泛化能力受限。为此,我们提出一种基于大规模网络爬取数据的生成式人脸语言模型预训练方法。利用互联网中包含人脸的文本与图像,进行自监督预训练,包括掩码图像/语言建模(MILM)和图像-文本匹配(ITM)任务。生成阶段进一步引入图像-文本匹配损失,使生成分布贴近控制信号,实现可控生成。实验表明,该模型在属性分类、表情识别等各类人脸下游任务中表现媲美当前最优预训练模型。此外,该方法还可广泛应用于人脸属性编辑、表情操控、遮挡物移除及图像修复等任务。

原文摘要 · Abstract (English)

Compared to the prosperity of pre-training models in natural image understanding, the research on large-scale pre-training models for facial knowledge learning is still limited. Current approaches mainly rely on manually assembled and annotated face datasets for training, but labeling such datasets is labor-intensive and the trained models have limited scalability beyond the training data. To address these limitations, we present a generative pre-training model for facial knowledge learning that leverages large-scale web-built data for training. We use texts and images containing human faces crawled from the internet and conduct pre-training on self-supervised tasks, including masked image/language modeling (MILM) and image-text matching (ITM). During the generation stage, we further utilize the image-text matching loss to pull the generation distribution towards the control signal for controllable image/text generation. Experimental results demonstrate that our model achieves comparable performance to state-of-the-art pre-training models for various facial downstream tasks, such as attribution classification and expression recognition. Furthermore, our approach is also applicable to a wide range of face editing tasks, including face attribute editing, expression manipulation, mask removal, and photo inpainting.

人脸生成自监督学习跨模态可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。