用指令微调让大模型看脸识情绪和特征,还能推理描述。
Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
- 基于人脸区域引导的注意力机制,融合几何与局部视觉特征。
- 在9个数据集5项任务中超越开源模型,接近商用水平。
- 生成描述更利于AI推理,适合社交智能与视觉语言研究。
人脸在社会交流中起核心作用,需高效计算机视觉工具支持以人为中心的应用。我们提出Face-LLaVA,一个面向人脸的多模态大语言模型,支持上下文学习,可进行表情与属性识别,并生成自然语言描述用于推理。我们首先构建了包含100万条样本的面部指令微调数据库FaceInstruct-1M;随后设计了一种基于人脸区域引导交叉注意力的新视觉编码器,整合面部几何结构与局部视觉特征。我们在九个不同数据集上评估了该方法,涵盖五项人脸处理任务:表情识别、动作单元检测、面部属性识别、年龄估计及深度伪造检测。结果表明,Face-LLaVA在开源多模态大模型中表现更优,且性能媲美商业方案。其输出在零样本设置下获得GPT更高的推理评分。相关数据集与模型将公开于https://face-llava.github.io,以推动社会人工智能与基础视觉语言研究的发展。
原文摘要 · Abstract (English)
The human face plays a central role in social communication, necessitating the use of performant computer vision tools for human-centered applications. We propose Face-LLaVA, a multimodal large language model for face-centered, in-context learning, including facial expression and attribute recognition. Additionally, Face-LLaVA is able to generate natural language descriptions that can be used for reasoning. Leveraging existing visual databases, we first developed FaceInstruct-1M, a face-centered database for instruction tuning MLLMs for face processing. We then developed a novel face-specific visual encoder powered by Face-Region Guided Cross-Attention that integrates face geometry with local visual features. We evaluated the proposed method across nine different datasets and five different face processing tasks, including facial expression recognition, action unit detection, facial attribute detection, age estimation and deepfake detection. Face-LLaVA achieves superior results compared to existing open-source MLLMs and competitive performance compared to commercial solutions. Our model output also receives a higher reasoning rating by GPT under a zero-shot setting across all the tasks. Both our dataset and model wil be released at https://face-llava.github.io to support future advancements in social AI and foundational vision-language research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。