arXiv:2410.03331cs.CV2024-10被引 10

在头戴设备遮挡面部情况下,实现表情识别的数据库与基础模型研究

EmojiHeroVR: A Study on Facial Expression Recognition under Partial Occlusion from Head-Mounted Displays

  • 构建含3556张标注图像的VR表情数据库,支持动态识别
  • 在遮挡条件下静态表情识别准确率达69.84%
  • 适合关注虚拟现实情感交互的研究者与开发者

情感识别通过提供情绪反馈和实现个性化,有助于提升虚拟现实(VR)体验。然而,由于头戴显示器(HMD)遮挡了面部上半部分,面部表情极少被用于识别用户情绪。为解决此问题,我们开展了包含37名参与者的新颖情感类VR游戏EmojiHeroVR实验。所收集的数据库EmoHeVRDB包含1,778个重新演绎情绪的3,556张标注面部图像。每张标注图像还附带前后各29帧视频,以支持动态面部表情识别(FER)。此外,每帧数据均包含通过Meta Quest Pro VR头显记录的63个面部动作单元激活信息。基于该数据库,我们使用EfficientNet-B0架构对六种基本情绪和中性情绪进行静态FER分类基准评估,最佳模型在测试集上达到69.84%的准确率,表明在HMD遮挡下进行表情识别虽具挑战性但可行。

原文摘要 · Abstract (English)

Emotion recognition promotes the evaluation and enhancement of Virtual Reality (VR) experiences by providing emotional feedback and enabling advanced personalization. However, facial expressions are rarely used to recognize users' emotions, as Head-Mounted Displays (HMDs) occlude the upper half of the face. To address this issue, we conducted a study with 37 participants who played our novel affective VR game EmojiHeroVR. The collected database, EmoHeVRDB (EmojiHeroVR Database), includes 3,556 labeled facial images of 1,778 reenacted emotions. For each labeled image, we also provide 29 additional frames recorded directly before and after the labeled image to facilitate dynamic Facial Expression Recognition (FER). Additionally, EmoHeVRDB includes data on the activations of 63 facial expressions captured via the Meta Quest Pro VR headset for each frame. Leveraging our database, we conducted a baseline evaluation on the static FER classification task with six basic emotions and neutral using the EfficientNet-B0 architecture. The best model achieved an accuracy of 69.84% on the test set, indicating that FER under HMD occlusion is feasible but significantly more challenging than conventional FER.

表情识别虚拟现实头部追踪数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。