arXiv:2507.10300cs.CVcs.AI2025-07ICCV被引 18

用合成数据训练的面部理解大模型,提升人脸细粒度分析能力

FaceLLM: A Multimodal Large Language Model for Face Understanding

  • 基于ChatGPT生成带属性提示的问答对构建训练集
  • 在表情、姿态、肤色纹理等任务上超越现有模型表现
  • 适合需要高精度人脸分析的研究与应用开发者

多模态大语言模型在视觉-语言任务中表现卓越,但现有模型主要基于通用数据集训练,难以处理人脸图像中的领域特定视觉线索。由于缺乏大规模标注的人脸图像-文本数据集,关于面部结构、表情、情绪及人口统计特征的细粒度理解仍不充分。本文提出FaceLLM,一个专用于人脸图像理解的多模态大语言模型。为构建训练数据,我们设计了一种新型弱监督流程,利用带有属性提示的ChatGPT生成来自FairFace数据集图像的高质量问答对,形成名为FairFaceGPT的数据集,涵盖表情、姿态、皮肤纹理和法医信息等多种属性。实验表明,FaceLLM在多项以人脸为中心的任务中显著优于现有模型,并达到当前最优水平。该工作展示了通过语言模型进行合成监督构建领域专用多模态模型的潜力,为可信、以人为本的多模态AI系统树立了范例。FairFaceGPT数据集与预训练的FaceLLM模型已在项目页面公开。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have shown remarkable performance in vision-language tasks. However, existing MLLMs are primarily trained on generic datasets, limiting their ability to reason on domain-specific visual cues such as those in facial images. In particular, tasks that require detailed understanding of facial structure, expression, emotion, and demographic features remain underexplored by MLLMs due to the lack of large-scale annotated face image-text datasets. In this work, we introduce FaceLLM, a multimodal large language model trained specifically for facial image understanding. To construct the training data, we propose a novel weakly supervised pipeline that uses ChatGPT with attribute-aware prompts to generate high-quality question-answer pairs based on images from the FairFace dataset. The resulting corpus, called FairFaceGPT, covers a diverse set of attributes including expression, pose, skin texture, and forensic information. Our experiments demonstrate that FaceLLM improves the performance of MLLMs on various face-centric tasks and achieves state-of-the-art performance. This work highlights the potential of synthetic supervision via language models for building domain-specialized MLLMs, and sets a precedent for trustworthy, human-centric multimodal AI systems. FairFaceGPT dataset and pretrained FaceLLM models are publicly available in the project page.

人脸理解多模态合成数据大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。