用合成数据提升头显摄像头表情识别效果
Enhancing Facial Expression Recognition in Head-Mounted Displays with Synthetic Data

- 从正面图像生成头显视角图像,解决数据稀缺问题
- 合成数据训练模型在真实场景中表现更优
- 适合研究混合现实表情识别的开发者
面部表情识别(FER)在使用头戴式显示器(HMD)的混合现实环境中对社交互动至关重要。然而,由于隐私顾虑和头显平台多样性,从头戴式摄像头(HMC)收集FER数据十分困难。现有FER数据集因视角独特而不直接适用。数据不足限制了基于神经网络的HMC FER方法发展。为此,我们提出一种数据合成框架,可将正面视角图像转换为HMC视角图像,利用现有大量标注数据。具体而言,先从图像重建3D纹理网格,再通过可配置相机系统渲染HMC视角图像。此外,引入纹理空间对齐网络(TSAN),实现精确纹理采样以保留细节表情。在模拟和真实HMC数据集上进行大量实验,结果表明:在合成数据上训练的模型优于现有数据训练模型,并在不同相机配置间表现出更好泛化能力。
原文摘要 · Abstract (English)
Facial expression recognition (FER) is crucial for social interaction in mixed reality environments that employ head-mounted displays (HMD). However, collecting FER data from head-mounted cameras (HMC) is challenging due to privacy concerns and the diversity of HMD platforms. Moreover, existing FER datasets are not directly applicable due to the unique perspectives of HMCs. The lack of sufficient data hinders the development of neural network-based HMC FER methods. To address data scarcity, we propose a data synthesis framework that generates HMC-view images from frontal-view images, leveraging abundant existing annotated datasets. Specifically, we first reconstruct 3D textured meshes from images and then apply a configurable camera system to render images from the HMC perspective. Additionally, we introduce a texture-space alignment network (TSAN) that enables accurate texture sampling from images to preserve detailed facial expressions. To evaluate the proposed method, we conduct extensive experiments on both simulated and real HMC datasets. Experimental results demonstrate that models trained on our synthetic dataset outperform those trained on existing datasets and exhibit better generalization across different camera configurations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。