arXiv:2604.09324cs.CV2026-04

用单目视频重建带精细表情和手势的逼真人体模型

Structure-Aware Fine-Grained Gaussian Splatting for Expressive Avatar Reconstruction

  • 结合时空三平面与时间感知六平面捕捉动态特征
  • 通过结构感知高斯模块提升姿态相关细节表现力
  • 专攻手部形变建模,支持单阶段训练生成自然动作

从单目视频中重建逼真且拓扑一致的人体虚拟形象仍是计算机视觉与图形学中的重大挑战。现有3D人体建模方法虽能有效捕捉身体运动,却难以准确建模手部动作与面部表情等细微特征。为此,我们提出结构感知细粒度高斯点渲染(SFGS),一种从单目视频序列重建具表现力且连贯的全身3D人体模型的新方法。SFGS利用仅空间的三平面与时间感知六平面,捕捉连续帧间的动态特征。设计结构感知高斯模块,在空间上保持一致性地捕捉依赖姿态的细节,提升姿态与纹理表达能力。为更好建模手部形变,还提出基于细粒度手部重建的残差精修模块。该方法仅需单阶段训练,在定量与定性评估中均优于现有最先进方法,生成具有自然运动与精细细节的高保真人体模型。代码已开源:https://github.com/Su245811YZ/SFGS

原文摘要 · Abstract (English)

Reconstructing photorealistic and topology-aware human avatars from monocular videos remains a significant challenge in the fields of computer vision and graphics. While existing 3D human avatar modeling approaches can effectively capture body motion, they often fail to accurately model fine details such as hand movements and facial expressions. To address this, we propose Structure-aware Fine-grained Gaussian Splatting (SFGS), a novel method for reconstructing expressive and coherent full-body 3D human avatars from a monocular video sequence. The SFGS use both spatial-only triplane and time-aware hexplane to capture dynamic features across consecutive frames. A structure-aware gaussian module is designed to capture pose-dependent details in a spatially coherent manner and improve pose and texture expression. To better model hand deformations, we also propose a residual refinement module based on fine-grained hand reconstruction. Our method requires only a single-stage training and outperforms state-of-the-art baselines in both quantitative and qualitative evaluations, generating high-fidelity avatars with natural motion and fine details. The code is on Github: https://github.com/Su245811YZ/SFGS

3D人体重建高斯点渲染手势建模单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。