让AI聚焦人脸任意区域,精准描述表情、动作和年龄。
Focal-RegionFace: Generating Fine-Grained Multi-attribute Descriptions for Arbitrarily Selected Face Focal Regions
- 分阶段微调视觉语言模型,逐步聚焦面部局部特征
- 在新数据集上实现表情、动作单元与年龄估计的高精度识别
- 适合需要细粒度人脸分析的研究与应用
本文提出一个未充分研究的问题:为任意选定的人脸区域生成并识别包含面部动作单元(AUs)、情绪状态和年龄估计的多属性自然语言描述(称为FaceFocalDesc)。我们认为,系统对特定面部区域的聚焦能力有助于提升理解与控制效果。为此,我们构建了一个面向任意面部区域的多属性描述数据集,提供丰富的区域级标注和自然语言描述。进一步,我们基于Qwen2.5-VL提出一种微调后的视觉语言模型Focal-RegionFace,通过多个渐进式微调阶段,逐步精细化关注局部面部特征,实现可解释的年龄估计、面部动作单元(FAU)和情绪检测。实验结果表明,Focal-RegionFace在新基准上各项传统与新提出指标均表现最佳,充分验证了其在细粒度多属性人脸区域聚焦分析场景中的有效性与通用性。
原文摘要 · Abstract (English)
In this paper, we introduce an underexplored problem in facial analysis: generating and recognizing multi-attribute natural language descriptions, containing facial action units (AUs), emotional states, and age estimation, for arbitrarily selected face regions (termed FaceFocalDesc). We argue that the system's ability to focus on individual facial areas leads to better understanding and control. To achieve this capability, we construct a new multi-attribute description dataset for arbitrarily selected face regions, providing rich region-level annotations and natural language descriptions. Further, we propose a fine-tuned vision-language model based on Qwen2.5-VL, called Focal-RegionFace for facial state analysis, which incrementally refines its focus on localized facial features through multiple progressively fine-tuning stages, resulting in interpretable age estimation, FAU and emotion detection. Experimental results show that Focal-RegionFace achieves the best performance on the new benchmark in terms of traditional and widely used metrics, as well as new proposed metrics. This fully verifies its effectiveness and versatility in fine-grained multi-attribute face region-focal analysis scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。