arXiv:2507.14608cs.CVcs.AI2025-07被引 2

用图模型捕捉面部特征间结构关系,提升表情识别准确率

Exp-Graph: How Connections Learn Facial Attributes in Graph-based Expression Recognition

  • 以关键点为节点,结合局部外观相似性构建面部特征图
  • 在三个数据集上分别达到98.09%、79.01%、56.39%准确率
  • 适合需要高精度表情识别的交互系统与医疗分析场景

面部表情识别在人脸动画、视频监控、情感计算和医学分析等人机交互应用中至关重要。由于面部特征结构随表情变化,融入结构信息对表情识别极为关键。本文提出Exp-Graph,一种基于图建模的新型框架,通过图结构表示面部特征间的关联关系。面部关键点作为图的顶点,边由关键点间距和视觉变压器编码的局部外观相似性决定。利用图卷积网络捕获并整合这些结构依赖,增强面部特征编码的表达能力。视觉变压器与图卷积模块协同,有效挖掘面部特征的局部与全局依赖,有助于表情识别。在Oulu-CASIA、eNTERFACE05和AFEW三个基准数据集上进行评估,准确率分别为98.09%、79.01%和56.39%。结果表明,Exp-Graph在受控实验环境与真实复杂场景下均具备强泛化能力,验证了其在实际应用中的有效性。

原文摘要 · Abstract (English)

Facial expression recognition is crucial for human-computer interaction applications such as face animation, video surveillance, affective computing, medical analysis, etc. Since the structure of facial attributes varies with facial expressions, incorporating structural information into facial attributes is essential for facial expression recognition. In this paper, we propose Exp-Graph, a novel framework designed to represent the structural relationships among facial attributes using graph-based modeling for facial expression recognition. For facial attributes graph representation, facial landmarks are used as the graph's vertices. At the same time, the edges are determined based on the proximity of the facial landmark and the similarity of the local appearance of the facial attributes encoded using the vision transformer. Additionally, graph convolutional networks are utilized to capture and integrate these structural dependencies into the encoding of facial attributes, thereby enhancing the accuracy of expression recognition. Thus, Exp-Graph learns from the facial attribute graphs highly expressive semantic representations. On the other hand, the vision transformer and graph convolutional blocks help the framework exploit the local and global dependencies among the facial attributes that are essential for the recognition of facial expressions. We conducted comprehensive evaluations of the proposed Exp-Graph model on three benchmark datasets: Oulu-CASIA, eNTERFACE05, and AFEW. The model achieved recognition accuracies of 98.09\%, 79.01\%, and 56.39\%, respectively. These results indicate that Exp-Graph maintains strong generalization capabilities across both controlled laboratory settings and real-world, unconstrained environments, underscoring its effectiveness for practical facial expression recognition applications.

表情识别图神经网络视觉变压器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。