arXiv:2508.13524cs.CVcs.AI2025-08被引 2

对比开源视觉语言模型与传统模型在人脸表情识别中的表现

Evaluating Open-Source Vision Language Models for Facial Emotion Recognition against Traditional Deep Learning Models

  • 用GFPGAN修复低质图像后评估模型性能
  • EfficientNet-B0准确率达86.44%,显著优于VLMs
  • 为情绪识别提供可复现的基准和部署成本分析

人脸表情识别(FER)在人机交互和心理健康诊断中至关重要。本研究首次在具有挑战性的FER-2013数据集上实证比较开源视觉语言模型(VLMs),包括Phi-3.5 Vision和CLIP,与传统深度学习模型VGG19、ResNet-50和EfficientNet-B0的表现。该数据集包含35,887张低分辨率灰度图像,涵盖七个情绪类别。为应对VLM训练假设与FER数据噪声之间的不匹配,我们提出一种新流程,结合基于GFPGAN的图像修复与FER评估。结果显示,传统模型表现显著优于VLMs:EfficientNet-B0达到86.44%准确率,ResNet-50为85.72%,而CLIP仅64.07%,Phi-3.5 Vision更低至51.66%。除精度、召回率、F1分数和准确率外,还详细分析了预处理、训练、推理和评估各阶段的计算开销,为实际部署提供参考。研究强调需改进VLM在噪声环境下的适应性,并为未来情绪识别研究提供可复现的基准。

原文摘要 · Abstract (English)

Facial Emotion Recognition (FER) is crucial for applications such as human-computer interaction and mental health diagnostics. This study presents the first empirical comparison of open-source Vision-Language Models (VLMs), including Phi-3.5 Vision and CLIP, against traditional deep learning models VGG19, ResNet-50, and EfficientNet-B0 on the challenging FER-2013 dataset, which contains 35,887 low-resolution grayscale images across seven emotion classes. To address the mismatch between VLM training assumptions and the noisy nature of FER data, we introduce a novel pipeline that integrates GFPGAN-based image restoration with FER evaluation. Results show that traditional models, particularly EfficientNet-B0 (86.44%) and ResNet-50 (85.72%), significantly outperform VLMs like CLIP (64.07%) and Phi-3.5 Vision (51.66%), highlighting the limitations of VLMs in low-quality visual tasks. In addition to performance evaluation using precision, recall, F1-score, and accuracy, we provide a detailed computational cost analysis covering preprocessing, training, inference, and evaluation phases, offering practical insights for deployment. This work underscores the need for adapting VLMs to noisy environments and provides a reproducible benchmark for future research in emotion recognition.

表情识别视觉语言模型模型对比图像修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。