用场景几何信息精准定位3D人体,提升动画真实感。
SHARE: Scene-Human Aligned Reconstruction

- 基于单目视频与场景点云对齐人体位置
- 关键帧优化+非关键帧相对位置保持,提升稳定性
- 适用于游戏、VR及野外视频,效果优于现有方法
在游戏、AR/VR和机器人领域,让角色自然互动环境至关重要。然而,当前人体运动重建难以准确放置人体于3D空间。本文提出场景-人体对齐重建(SHARE),利用场景几何提供的空间线索,实现人体运动的精准定位。仅需固定摄像头的单目RGB视频,共享首先逐帧估计人体网格与分割掩码,并在关键帧生成场景点图。通过将人体网格与从场景中提取的人体点云对比,迭代优化关键帧中的人体位置。关键创新在于:在优化过程中,保持非关键帧人体网格相对于关键帧根关节的相对位置一致性。该方法在重建场景的同时实现更精确的人体定位,适用于精心构建的数据集和真实网络视频。大量实验表明,SHARE显著优于现有方法。
原文摘要 · Abstract (English)
Animating realistic character interactions with the surrounding environment is important for autonomous agents in gaming, AR/VR, and robotics. However, current methods for human motion reconstruction struggle with accurately placing humans in 3D space. We introduce Scene-Human Aligned REconstruction (SHARE), a technique that leverages the scene geometry's inherent spatial cues to accurately ground human motion reconstruction. Each reconstruction relies solely on a monocular RGB video from a stationary camera. SHARE first estimates a human mesh and segmentation mask for every frame, alongside a scene point map at keyframes. It iteratively refines the human's positions at these keyframes by comparing the human mesh against the human point map extracted from the scene using the mask. Crucially, we also ensure that non-keyframe human meshes remain consistent by preserving their relative root joint positions to keyframe root joints during optimization. Our approach enables more accurate 3D human placement while reconstructing the surrounding scene, facilitating use cases on both curated datasets and in-the-wild web videos. Extensive experiments demonstrate that SHARE outperforms existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。