用单个摄像头实现逼真全身虚拟化身,实时驱动且贴合视角。
EgoAvatar: Egocentric View-Driven and Photorealistic Full-body Avatars
- 基于骨骼运动驱动可动画化角色模型,兼顾几何与外观建模。
- 从单眼视频恢复全身动作,实现个性化动作捕捉。
- 测试时网格精修确保形象精准投影到第一人称画面。
沉浸式虚拟现实远程呈现的理想状态是能够与数字化身交互,使其在行为上与真实人物无异。核心挑战在于:构建能精确反映真实人类的数字孪生,并仅通过轻量、低功耗的单眼传感设备(如单个RGB相机)进行人体追踪。现有工作或仅关注第一人称动作捕捉,或仅建模头部,或依赖多视角采集。本文首次提出一种面向特定个体的第一人称远程呈现方法,可联合建模逼真的数字化身并由单个第一人称视频驱动。首先构建一个可由骨骼运动驱动、同时支持几何与外观建模的角色模型;其次引入个性化第一人称动作捕捉模块,从第一人称视频中恢复全身姿态;最后将恢复的姿态应用于角色模型,并在测试阶段进行网格精修,使几何结构准确投影至第一人称视图。为验证设计选择,我们构建了一个新的挑战性基准,包含真人执行多种动作时的第一人称与密集多视角视频对。实验表明,该方法显著优于基线及竞争方法,朝着第一人称与逼真远程呈现迈出关键一步。
原文摘要 · Abstract (English)
Immersive VR telepresence ideally means being able to interact and communicate with digital avatars that are indistinguishable from and precisely reflect the behaviour of their real counterparts. The core technical challenge is two fold: Creating a digital double that faithfully reflects the real human and tracking the real human solely from egocentric sensing devices that are lightweight and have a low energy consumption, e.g. a single RGB camera. Up to date, no unified solution to this problem exists as recent works solely focus on egocentric motion capture, only model the head, or build avatars from multi-view captures. In this work, we, for the first time in literature, propose a person-specific egocentric telepresence approach, which jointly models the photoreal digital avatar while also driving it from a single egocentric video. We first present a character model that is animatible, i.e. can be solely driven by skeletal motion, while being capable of modeling geometry and appearance. Then, we introduce a personalized egocentric motion capture component, which recovers full-body motion from an egocentric video. Finally, we apply the recovered pose to our character model and perform a test-time mesh refinement such that the geometry faithfully projects onto the egocentric view. To validate our design choices, we propose a new and challenging benchmark, which provides paired egocentric and dense multi-view videos of real humans performing various motions. Our experiments demonstrate a clear step towards egocentric and photoreal telepresence as our method outperforms baselines as well as competing methods. For more details, code, and data, we refer to our project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。