arXiv:2601.03507cs.CV2026-01CVPR被引 7

用VR眼镜摄像头实时捕捉面部表情,无需校准即可驱动虚拟形象。

REFA: Real-time Egocentric Facial Animations for Virtual Reality

  • 基于多源数据蒸馏训练模型,融合合成与真实图像。
  • 采集1.8万组用户数据,通过可微渲染自动标注表情。
  • 适合虚拟会议、游戏等需要自然表情交互的场景。

我们提出一种新型系统,利用嵌入在虚拟现实(VR)头显中的红外摄像头捕获第一人称视角,实现实时面部表情追踪。该技术使用户无需繁琐校准即可非侵入式地精准驱动虚拟角色的面部表情。系统核心为基于知识蒸馏的方法,可在异构数据和标签来源(如合成图像与真实图像)上训练机器学习模型。作为数据集的一部分,我们使用仅含手机和定制带额外摄像头的轻量级采集设备,收集了1.8万组多样化用户的面部数据。为处理这些数据,我们开发了一套鲁棒的可微分渲染流水线,实现面部表情标签的自动提取。本系统为虚拟环境中的人际交流与表达开辟了新路径,适用于视频会议、游戏、娱乐及远程协作等应用。

原文摘要 · Abstract (English)

We present a novel system for real-time tracking of facial expressions using egocentric views captured from a set of infrared cameras embedded in a virtual reality (VR) headset. Our technology facilitates any user to accurately drive the facial expressions of virtual characters in a non-intrusive manner and without the need of a lengthy calibration step. At the core of our system is a distillation based approach to train a machine learning model on heterogeneous data and labels coming form multiple sources, \eg synthetic and real images. As part of our dataset, we collected 18k diverse subjects using a lightweight capture setup consisting of a mobile phone and a custom VR headset with extra cameras. To process this data, we developed a robust differentiable rendering pipeline enabling us to automatically extract facial expression labels. Our system opens up new avenues for communication and expression in virtual environments, with applications in video conferencing, gaming, entertainment, and remote collaboration.

虚拟现实表情捕捉实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。