arXiv:2505.09336cs.CV2025-05被引 1

用伪标签引导跨视角面部表情学习,无监督提升识别精度

Unsupervised Multiview Contrastive Language-Image Joint Learning with Pseudo-Labeled Prompts Via Vision-Language Model for 3D/4D Facial Expression Recognition

  • 用生成文本提示生成伪标签,指导情绪语义对齐
  • 跨视角联合嵌入空间提升特征一致性,准确率超现有方法
  • 适合缺乏标注数据的实时表情识别场景

本文提出MultiviewVLM,一种用于3D/4D面部表情识别的无监督多视角视觉-语言联合表示学习模型。通过生成的文本提示提取伪标签,引导情绪语义的隐式对齐。为捕捉多视角共享信息,设计了无需显式监督的联合嵌入空间。引入一种稳定的正负样本采样策略,增强模型判别能力;同时采用梯度友好的损失函数,实现更平滑稳定的收敛,并支持分布式训练以保证可扩展性。大量实验表明,MultiviewVLM超越现有最优方法,且只需少量修改即可适配多种实际应用。

原文摘要 · Abstract (English)

In this paper, we introduce MultiviewVLM, a vision-language model designed for unsupervised contrastive multiview representation learning of facial emotions from 3D/4D data. Our architecture integrates pseudo-labels derived from generated textual prompts to guide implicit alignment of emotional semantics. To capture shared information across multi-views, we propose a joint embedding space that aligns multiview representations without requiring explicit supervision. We further enhance the discriminability of our model through a novel multiview contrastive learning strategy that leverages stable positive-negative pair sampling. A gradient-friendly loss function is introduced to promote smoother and more stable convergence, and the model is optimized for distributed training to ensure scalability. Extensive experiments demonstrate that MultiviewVLM outperforms existing state-of-the-art methods and can be easily adapted to various real-world applications with minimal modifications.

表情识别多视角学习无监督视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。