arXiv:2507.05620cs.CVcs.LG2025-07

用无配对头戴摄像头数据生成逼真虚拟人像,无需昂贵同步采集。

Generative Head-Mounted Camera Captures for Photorealistic Avatars

  • 利用大量易获取的无配对头戴相机数据,直接生成高质量合成图像。
  • 成功分离表情与外观,实现更精准的虚拟人姿态控制。
  • 可泛化到未见身份,减少对成对数据依赖,适合大规模应用。

在虚拟现实与增强现实环境中实现逼真虚拟人像动画面临挑战,主要源于难以获取面部真实状态。头戴式摄像机(HMC)因物理限制无法同步采集红外与外部全景摄像机的完整视角数据,而现有依赖分析-合成方法虽能生成准确真实值,但个性化训练中表情与风格解耦不彻底。此外,需为同一主体收集大量成对(HMC与全景)数据,操作成本高,且无法跨视角、光照复用。本文提出新型生成方法GenHMC,利用大规模无配对的HMC数据,在给定全景捕捉的虚拟人姿态条件下,直接生成高质量合成HMC图像。实验表明,该方法能有效分离输入条件中的表情与视角信号,独立于面部外观,提升真实度;并具备跨身份泛化能力,摆脱成对数据依赖。通过评估合成图像及由此训练的通用人脸编码器,验证了其更高的数据效率和当前最优精度。

原文摘要 · Abstract (English)

Enabling photorealistic avatar animations in virtual and augmented reality (VR/AR) has been challenging because of the difficulty of obtaining ground truth state of faces. It is physically impossible to obtain synchronized images from head-mounted cameras (HMC) sensing input, which has partial observations in infrared (IR), and an array of outside-in dome cameras, which have full observations that match avatars' appearance. Prior works relying on analysis-by-synthesis methods could generate accurate ground truth, but suffer from imperfect disentanglement between expression and style in their personalized training. The reliance of extensive paired captures (HMC and dome) for the same subject makes it operationally expensive to collect large-scale datasets, which cannot be reused for different HMC viewpoints and lighting. In this work, we propose a novel generative approach, Generative HMC (GenHMC), that leverages large unpaired HMC captures, which are much easier to collect, to directly generate high-quality synthetic HMC images given any conditioning avatar state from dome captures. We show that our method is able to properly disentangle the input conditioning signal that specifies facial expression and viewpoint, from facial appearance, leading to more accurate ground truth. Furthermore, our method can generalize to unseen identities, removing the reliance on the paired captures. We demonstrate these breakthroughs by both evaluating synthetic HMC images and universal face encoders trained from these new HMC-avatar correspondences, which achieve better data efficiency and state-of-the-art accuracy.

虚拟人像生成模型头戴相机姿态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。