用VR头显捕捉面部激活数据,提升虚拟现实中的表情识别准确率
Unimodal and Multimodal Static Facial Expression Recognition for Virtual Reality Users with EmoHeVRDB
- 基于Meta Quest Pro的面部激活数据进行静态表情识别
- 多模态融合达80.42%准确率,显著高于图像基线的69.84%
- 首次在EmoHeVRDB上实现单模与多模态表情识别新基准
本研究探索了利用Meta Quest Pro VR头显捕获的面部表情激活(FEA)数据,在虚拟现实环境中进行静态面部表情识别(FER)的可行性。基于EmojiHeroVR数据库(EmoHeVRDB),我们对比了多种单模态方法,在七类情绪识别任务中最高达到73.02%的准确率。进一步融合FEA与图像数据的多模态方法显著提升性能,中间融合策略取得80.42%最高准确率,显著优于EmoHeVRDB图像数据基准的69.84%。本研究是首个利用EmoHeVRDB独特FEA数据开展单模与多模态静态FER的工作,为VR场景下的表情识别建立了新基准。结果表明,融合互补模态可有效克服头戴式显示器(HMD)导致的遮挡问题,显著提升识别精度。
原文摘要 · Abstract (English)
In this study, we explored the potential of utilizing Facial Expression Activations (FEAs) captured via the Meta Quest Pro Virtual Reality (VR) headset for Facial Expression Recognition (FER) in VR settings. Leveraging the EmojiHeroVR Database (EmoHeVRDB), we compared several unimodal approaches and achieved up to 73.02% accuracy for the static FER task with seven emotion categories. Furthermore, we integrated FEA and image data in multimodal approaches, observing significant improvements in recognition accuracy. An intermediate fusion approach achieved the highest accuracy of 80.42%, significantly surpassing the baseline evaluation result of 69.84% reported for EmoHeVRDB's image data. Our study is the first to utilize EmoHeVRDB's unique FEA data for unimodal and multimodal static FER, establishing new benchmarks for FER in VR settings. Our findings highlight the potential of fusing complementary modalities to enhance FER accuracy in VR settings, where conventional image-based methods are severely limited by the occlusion caused by Head-Mounted Displays (HMDs).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。