用视觉语言模型实现3D/4D表情识别,无需标注也能达到顶尖效果
Self-Supervised Multi-View Representation Learning using Vision-Language Model for 3D/4D Facial Expression Recognition
- 通过多视角去相关+跨模态对齐,学习鲁棒且不变的面部表征
- 在多个数据集上超越无监督方法,媲美甚至超过有监督基线
- 适合缺乏标注数据的实时情绪分析与微表情识别场景
表情识别是情感计算的基础任务,广泛应用于人机交互、心理健康分析和行为理解。本文提出SMILE-VLM,一种用于3D/4D表情识别的自监督视觉语言模型,统一多视角视觉表征学习与自然语言监督。该模型通过三个核心组件:基于Barlow Twins风格的多视角去相关、视觉-语言对比对齐、跨模态冗余最小化,学习具有语义一致性且视角不变的嵌入表示。实验表明,该框架在多个基准上达到当前最优性能。进一步扩展至4D微表情识别任务,可捕捉细微情感线索。大量结果证明,SMILE-VLM不仅显著优于现有无监督方法,还匹配或超越有监督基线,为表达性面部行为理解提供了一种可扩展、低标注依赖的解决方案。
原文摘要 · Abstract (English)
Facial expression recognition (FER) is a fundamental task in affective computing with applications in human-computer interaction, mental health analysis, and behavioral understanding. In this paper, we propose SMILE-VLM, a self-supervised vision-language model for 3D/4D FER that unifies multiview visual representation learning with natural language supervision. SMILE-VLM learns robust, semantically aligned, and view-invariant embeddings by proposing three core components: multiview decorrelation via a Barlow Twins-style loss, vision-language contrastive alignment, and cross-modal redundancy minimization. Our framework achieves the state-of-the-art performance on multiple benchmarks. We further extend SMILE-VLM to the task of 4D micro-expression recognition (MER) to recognize the subtle affective cues. The extensive results demonstrate that SMILE-VLM not only surpasses existing unsupervised methods but also matches or exceeds supervised baselines, offering a scalable and annotation-efficient solution for expressive facial behavior understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。