arXiv:2512.06377cs.CV2025-12被引 2

为教育场景设计多维表情识别模型,首次在FER2013上标注主导维度并提升预测精度。

VAD-Net: Multidimensional Facial Expression Recognition in Intelligent Education System

  • 在FER2013上新增主导度(Dominance)标注,构建首个包含VAD三维度的标注数据集
  • 引入正交卷积增强网络表达能力,使三维度情绪预测准确率显著提升
  • 适合情感计算、智能教育系统开发者及多维情绪研究者使用

当前面部表情识别(FER)数据集多采用单一情绪类别标签(如快乐、愤怒、悲伤等),表达能力有限。未来情感计算需更全面精确的多维情绪度量,如通过效价-唤醒-主导度(VAD)三维度参数衡量。尽管AffectNet已添加效价(Valence)与唤醒(Arousal)信息,仍缺乏主导度(Dominance)。本研究首次在FER2013数据集上补充主导度标注,并提出基于ResNet的正交化回归网络以增强特征提取能力。实验表明,主导度虽可测量但人工标注和网络预测难度均高于效价与唤醒;引入正交卷积后,多维预测性能显著提升。所构建的带VAD标注的FER2013数据集可作为多维情绪评估基准,正交化网络亦可作为该任务基线模型。数据集与代码已公开于https://github.com/YeeHoran/VAD-Net。

原文摘要 · Abstract (English)

Current FER (Facial Expression Recognition) dataset is mostly labeled by emotion categories, such as happy, angry, sad, fear, disgust, surprise, and neutral which are limited in expressiveness. However, future affective computing requires more comprehensive and precise emotion metrics which could be measured by VAD(Valence-Arousal-Dominance) multidimension parameters. To address this, AffectNet has tried to add VA (Valence and Arousal) information, but still lacks D(Dominance). Thus, the research introduces VAD annotation on FER2013 dataset, takes the initiative to label D(Dominance) dimension. Then, to further improve network capacity, it enforces orthogonalized convolution on it, which extracts more diverse and expressive features and will finally increase the prediction accuracy. Experiment results show that D dimension could be measured but is difficult to obtain compared with V and A dimension no matter in manual annotation or regression network prediction. Secondly, the ablation test by introducing orthogonal convolution verifies that better VAD prediction could be obtained in the configuration of orthogonal convolution. Therefore, the research provides an initiative labelling for D dimension on FER dataset, and proposes a better prediction network for VAD prediction through orthogonal convolution. The newly built VAD annotated FER2013 dataset could act as a benchmark to measure VAD multidimensional emotions, while the orthogonalized regression network based on ResNet could act as the facial expression recognition baseline for VAD emotion prediction. The newly labeled dataset and implementation code is publicly available on https://github.com/YeeHoran/VAD-Net .

表情识别多维情绪教育科技正交卷积

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。