专攻人脸感知的多模态大模型,性能显著超越现有方法。
Face-MLLM: A Large Face Perception Model
- 构建细粒度人脸描述数据集,重构传统人脸数据为问答形式。
- 提出三阶段训练法,分步提升视觉-文本对齐与人脸任务理解能力。
- 在5项经典任务和新零样本属性分析任务中均表现优异,适合人脸分析研究者。
尽管多模态大语言模型(MLLMs)在众多视觉-语言任务上表现优异,但其对人脸的感知与理解能力尚未得到充分探索。本文系统评估了现有MLLMs在人脸感知任务上的表现,定量结果显示其普遍表现不佳,主因是缺乏包含细粒度人脸描述的图文数据集。为此,我们设计了一套实用的数据集构建流程:重新标注LAION-Face数据集,添加更详细的人脸描述与面部属性标签;并将传统人脸数据集重构为适合MLLMs的问答格式。基于这些增强数据,我们提出一种新型三阶段训练方法:第一阶段学习视觉-文本对齐,第二阶段掌握基础视觉问答能力,第三阶段专注多种专业人脸感知任务。实验表明,本模型在5个知名人脸感知任务中优于先前MLLMs,且在新提出的零样本面部属性分析任务中也表现卓越。
原文摘要 · Abstract (English)
Although multimodal large language models (MLLMs) have achieved promising results on a wide range of vision-language tasks, their ability to perceive and understand human faces is rarely explored. In this work, we comprehensively evaluate existing MLLMs on face perception tasks. The quantitative results reveal that existing MLLMs struggle to handle these tasks. The primary reason is the lack of image-text datasets that contain fine-grained descriptions of human faces. To tackle this problem, we design a practical pipeline for constructing datasets, upon which we further build a novel multimodal large face perception model, namely Face-MLLM. Specifically, we re-annotate LAION-Face dataset with more detailed face captions and facial attribute labels. Besides, we re-formulate traditional face datasets using the question-answer style, which is fit for MLLMs. Together with these enriched datasets, we develop a novel three-stage MLLM training method. In the first two stages, our model learns visual-text alignment and basic visual question answering capability, respectively. In the third stage, our model learns to handle multiple specialized face perception tasks. Experimental results show that our model surpasses previous MLLMs on five famous face perception tasks. Besides, on our newly introduced zero-shot facial attribute analysis task, our Face-MLLM also presents superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。