用头戴设备实现真人动作捕捉与光照还原,生成可随意重打光的逼真虚拟人。
EgoRelight: Egocentric Human Capture and Illumination Recovery for Relightable and Photoreal Avatar Rendering

- 通过头显下视双目摄像头获取稠密深度图,驱动网格化虚拟人模型。
- 神经外观模型分离处理镜面与漫反射光照,无需预设材质模型即可泛化到新光照。
- 实时逆渲染恢复环境光照图,让虚拟人自然融入真实场景,适合远程会议等应用。
混合现实头显有望实现沉浸式远程共在,使虚拟人无缝融入真实或虚拟环境。这需要从头戴设备受限视角同时完成用户动作捕捉、新光照下外观估计及环境理解。现有方法将这些问题割裂处理:要么依赖预烘焙光照驱动虚拟人,要么需在摄影棚内进行重打光。本文提出EgoRelight,一个统一的自指视角远程共在框架,能同步捕获全身动作、合成逼真且可重打光的外观,并从单个头显数据恢复高动态范围(HDR)环境图。首先,设计自指感知模块,利用下视双目相机提取稠密深度图,作为几何控制信号驱动基于网格的虚拟人模型。其次,提出新型神经外观模型,分别学习视点相关的镜面反射和视点无关的漫反射光照;通过专用射线采样策略,在不依赖限制性解析BRDF先验的情况下,泛化至未见光照条件。第三,通过测试时逆渲染过程,匹配预训练虚拟人外观与实时自指摄像头观测,恢复HDR环境图,实现虚拟人与物理世界的无缝融合。我们在社交远程共在场景中验证系统,远程用户能根据其真实环境一致地被重打光。大量实验表明,各组件及集成系统在几何精度、渲染质量与重打光保真度上均显著优于现有最先进方法。
原文摘要 · Abstract (English)
Mixed Reality (MR) headsets promise a future of immersive telepresence where virtual humans blend indistinguishably into real or virtual surroundings. Achieving this vision requires a method for capturing a user's motion, estimating appearance under novel lighting, and understanding the environment - all from the constrained viewpoint of a head-mounted display (HMD). Existing approaches treat these as isolated problems: they either focus on driving avatars with baked-in lighting or rely on studio setups for relighting. In this paper, we present EgoRelight, a holistic framework for egocentric telepresence that simultaneously captures full-body human performance, synthesizes photorealistic and relightable appearance, and estimates high dynamic range (HDR) environment maps from a single HMD. First, to ensure motion and surface reconstruction, we propose an egocentric perception module that leverages stereo down-facing cameras to extract dense depth maps, which serve as geometric control signals to drive a mesh-based avatar. Second, we introduce a novel neural appearance model that learns to synthesize view-dependent specular and view-independent diffuse shading separately. By employing a specialized ray-sampling strategy, our model generalizes to unseen illumination without relying on restrictive analytical BRDF priors. Third, we enable seamless avatar integration into the physical world via a test-time inverse rendering process, which recovers an HDR environment map by matching the pre-trained avatar's appearance to live egocentric camera observations. We demonstrate our system through a social telepresence application, where remote users are coherently relit according to their physical environment. Extensive experiments show that our components and the integrated system significantly outperform state-of-the-art baselines in geometric accuracy and rendering as well as relighting fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。