arXiv:2504.15624cs.CV2025-04被引 7

专为识脸设计的多模态大模型,能精准理解人脸细节。

FaceInsight: A Multimodal Large Language Model for Face Perception

  • 用视觉-文本对齐建模人脸信息的确定与不确定关系。
  • 引入人脸分割图作为辅助输入,提升局部结构理解能力。
  • 在三类识脸任务中均优于9个对比模型,无需微调也表现优异。

近年来,多模态大语言模型在通用视觉内容理解方面展现出强大能力,但在人脸感知任务中表现不佳,常对人脸相关问题给出不准确或误导性回答。为此,我们提出FaceInsight,一种面向人脸感知的多功能多模态大语言模型,可提供细粒度的面部信息。该方法通过建立面部知识的视觉-文本对齐,建模面部信息间的不确定依赖与确定性关系,缓解纯语言推理的局限性。此外,引入人脸分割图作为辅助感知模态,以局部结构线索丰富视觉输入,增强语义理解。在三个面部感知任务上的综合实验与分析表明,无论是否微调,FaceInsight均持续优于九种对比模型。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing inaccurate or misleading responses to face-specific queries. To address this gap, we propose FaceInsight, the versatile face perception MLLM that provides fine-grained facial information. Our approach introduces visual-textual alignment of facial knowledge to model both uncertain dependencies and deterministic relationships among facial information, mitigating the limitations of language-driven reasoning. Additionally, we incorporate face segmentation maps as an auxiliary perceptual modality, enriching the visual input with localized structural cues to enhance semantic understanding. Comprehensive experiments and analyses across three face perception tasks demonstrate that FaceInsight consistently outperforms nine compared MLLMs under both training-free and fine-tuned settings.

多模态人脸理解大模型视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。