用语言提示融合多视角面部表情,提升3D/4D表情识别准确率
Facial Emotion Learning with Text-Guided Multiview Fusion via Vision-Language Model for 3D/4D Facial Expression Recognition
- 通过语言引导的多视角特征融合,实现情绪语义对齐
- 在多个基准上达到当前最佳性能,包括BU-3DFE和BP4D-Spontaneous
- 适用于真实场景下的微表情识别,适合人机交互与医疗监测应用
3D/4D面部表情识别(FER)在情感计算中面临空间与时间动态复杂的挑战,对行为理解、健康监测和人机交互具有重要意义。本文提出FACET-VLM,一种基于视觉-语言模型的3D/4D FER框架,结合多视角面部表征学习与自然语言提示的语义引导。该框架包含三个核心组件:跨视角语义聚合(CVSA)实现视角一致性融合,多视角文本引导融合(MTGF)实现情绪语义对齐,以及多视角一致性损失以保证结构连贯性。模型在多个基准(包括BU-3DFE、Bosphorus、BU-4DFE和BP4D-Spontaneous)上达到当前最优性能。进一步扩展至4D微表情识别(MER)任务,在4DME数据集上展现出捕捉细微短暂情绪线索的能力。实验验证了各组件的有效性。总体而言,FACET-VLM为姿态与自发场景下的多模态表情识别提供了鲁棒、可扩展且高性能的解决方案。
原文摘要 · Abstract (English)
Facial expression recognition (FER) in 3D and 4D domains presents a significant challenge in affective computing due to the complexity of spatial and temporal facial dynamics. Its success is crucial for advancing applications in human behavior understanding, healthcare monitoring, and human-computer interaction. In this work, we propose FACET-VLM, a vision-language framework for 3D/4D FER that integrates multiview facial representation learning with semantic guidance from natural language prompts. FACET-VLM introduces three key components: Cross-View Semantic Aggregation (CVSA) for view-consistent fusion, Multiview Text-Guided Fusion (MTGF) for semantically aligned facial emotions, and a multiview consistency loss to enforce structural coherence across views. Our model achieves state-of-the-art accuracy across multiple benchmarks, including BU-3DFE, Bosphorus, BU-4DFE, and BP4D-Spontaneous. We further extend FACET-VLM to 4D micro-expression recognition (MER) on the 4DME dataset, demonstrating strong performance in capturing subtle, short-lived emotional cues. The extensive experimental results confirm the effectiveness and substantial contributions of each individual component within the framework. Overall, FACET-VLM offers a robust, extensible, and high-performing solution for multimodal FER in both posed and spontaneous settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。