用单张第一视角图生成可动虚拟人像,提升远程沉浸感
EgoAnimate: Generating Human Animations from Egocentric top-down Views
- 基于稳定扩散与ControlNet,从顶部视角图重建正面视图
- 仅需一张第一人称图像即可生成逼真可动虚拟人
- 无需多视角训练数据,通用性更强,适合轻量级应用
理想的数字远程存在体验需要精准还原人的身体、衣着和动作。采用第一人称(即第一视角)视角可借助便携低成本设备实现,无需前视摄像头,但该视角带来遮挡和身体比例失真的挑战。现有方法极少从第一视角重建人体外观,且均未使用生成先验方法。部分方法在推理时仅用单张第一视角图像生成虚拟人,但仍依赖多视角数据集进行训练。本文首次提出基于生成式骨干网络的方案,从第一视角输入重建可动画化虚拟人。基于Stable Diffusion架构,结合ControlNet,设计流程将被遮挡的顶部视角图像转化为真实感强的正面视图,再输入图像到动作模型中生成动作。该方法显著降低训练负担,提升泛化能力,仅需单张输入即可实现高质量虚拟人动画生成,为更普及、更通用的远程存在系统铺平道路。
原文摘要 · Abstract (English)
An ideal digital telepresence experience requires accurate replication of a person's body, clothing, and movements. To capture and transfer these movements into virtual reality, the egocentric (first-person) perspective can be adopted, which enables the use of a portable and cost-effective device without front-view cameras. However, this viewpoint introduces challenges such as occlusions and distorted body proportions. There are few works reconstructing human appearance from egocentric views, and none use a generative prior-based approach. Some methods create avatars from a single egocentric image during inference, but still rely on multi-view datasets during training. To our knowledge, this is the first study using a generative backbone to reconstruct animatable avatars from egocentric inputs. Based on Stable Diffusion, our method reduces training burden and improves generalizability. Inspired by methods such as SiTH and MagicMan, which perform 360-degree reconstruction from a frontal image, we introduce a pipeline that generates realistic frontal views from occluded top-down images using ControlNet and a Stable Diffusion backbone. Our goal is to convert a single top-down egocentric image into a realistic frontal representation and feed it into an image-to-motion model. This enables generation of avatar motions from minimal input, paving the way for more accessible and generalizable telepresence systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。