arXiv:2505.05043cs.CV2025-05

xTrace实时分析真实人脸表情,精准预测情绪的愉悦与唤醒度。

xTrace: A Facial Expressive Behaviour Analysis Tool for Continuous Affect Recognition

  • 基于45万条真实视频数据,覆盖广泛情绪空间,训练出高泛化能力模型。
  • 采用可解释且高效特征提取,实现低计算成本下的高精度情绪识别。
  • 在真实场景测试中性能超越现有方法7.1%,适合情绪研究与应用开发。

在情感计算领域,识别面部视频中的表达行为是一项长期挑战。尽管近年进展显著,但在真实自然场景下实现实时、鲁棒的面部表情分析系统仍具难度。本文针对两大关键挑战:一是缺乏大规模标注的面部情感视频数据集,覆盖维度情感空间不全;二是难以提取具有判别性、可解释性、鲁棒性且计算高效的面部特征。为此,本文提出xTrace,一个用于连续情绪识别的面部表达行为分析工具,可从真实环境人脸视频中预测维度情绪(愉悦度与唤醒度)。为解决第一项挑战,模型在目前最大规模的情感视频数据集上训练,包含约45万条视频,覆盖维度情感空间多数区域,使xTrace具备分析多样化自然表情的能力。为应对第二项挑战,xTrace采用可解释且高效的面部情感描述符,在保持低计算复杂度的同时实现高准确率与鲁棒性。关键组件在三个现有工具(MediaPipe、OpenFace、Augsburg Affect Toolbox)上进行对比评估。在包含约5万条视频的真实环境基准数据集上,xTrace达到0.86的平均一致性相关系数(CCC);在SEWA测试集上达0.75平均CCC,优于现有最先进方法约7.1%。

原文摘要 · Abstract (English)

Recognising expressive behaviours in face videos is a long-standing challenge in Affective Computing. Despite significant advancements in recent years, it still remains a challenge to build a robust and reliable system for naturalistic and in-the-wild facial expressive behaviour analysis in real time. This paper addresses two key challenges in building such a system: (1). The paucity of large-scale labelled facial affect video datasets with extensive coverage of the 2D emotion space, and (2). The difficulty of extracting facial video features that are discriminative, interpretable, robust, and computationally efficient. Toward addressing these challenges, this work introduces xTrace, a robust tool for facial expressive behaviour analysis and predicting continuous values of dimensional emotions, namely valence and arousal, from in-the-wild face videos. To address challenge (1), the proposed affect recognition model is trained on the largest facial affect video data set, containing $\sim$450k videos that cover most emotion zones in the dimensional emotion space, making xTrace highly versatile in analysing a wide spectrum of naturalistic expressive behaviours. To address challenge (2), xTrace uses facial affect descriptors that are not only explainable, but can also achieve a high degree of accuracy and robustness with low computational complexity. The key components of xTrace are benchmarked against three existing tools: MediaPipe, OpenFace, and Augsburg Affect Toolbox. On an in-the-wild benchmarking set composed of $\sim$50k videos, xTrace achieves 0.86 mean Concordance Correlation Coefficient (CCC) and on the SEWA test set it achieves 0.75 mean CCC, outperforming existing SOTA by $\sim$7.1\%.

情绪识别面部分析连续情绪真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。