用人体结构先验提升单目视频中动态人物重建的几何准确性
HuPrior3R: Incorporating Human Priors for Better 3D Dynamic Reconstruction from Monocular Videos
- 结合SMPL人体模型与单目深度估计,引入混合几何先验
- 在TUM Dynamics和GTA-IM上实现更准确的人体比例与边界保持
- 适合关注高精度3D人体重建的研究者与应用开发者
单目动态视频重建在复杂人体场景中面临几何不一致与分辨率下降问题。现有方法缺乏对3D人体结构的理解,导致肢体比例失真、人物与物体融合异常,且受内存限制的下采样使人体边界向背景几何漂移。为此,本文提出融合SMPL人体模型与单目深度估计的混合几何先验,利用结构化人体先验保持表面一致性并捕捉人体区域的细粒度几何细节。我们设计了层级化处理流程:先对全分辨率图像进行整体场景几何建模,再通过策略性裁剪与交叉注意力融合增强人体局部细节。特征融合模块整合SMPL先验,确保几何合理性的同时保留精细的人体边界。在TUM Dynamics与GTA-IM数据集上的大量实验表明,该方法在动态人体重建任务中表现更优。
原文摘要 · Abstract (English)
Monocular dynamic video reconstruction faces significant challenges in dynamic human scenes due to geometric inconsistencies and resolution degradation issues. Existing methods lack 3D human structural understanding, producing geometrically inconsistent results with distorted limb proportions and unnatural human-object fusion, while memory-constrained downsampling causes human boundary drift toward background geometry. To address these limitations, we propose to incorporate hybrid geometric priors that combine SMPL human body models with monocular depth estimation. Our approach leverages structured human priors to maintain surface consistency while capturing fine-grained geometric details in human regions. We introduce HuPrior3R, featuring a hierarchical pipeline with refinement components that processes full-resolution images for overall scene geometry, then applies strategic cropping and cross-attention fusion for human-specific detail enhancement. The method integrates SMPL priors through a Feature Fusion Module to ensure geometrically plausible reconstruction while preserving fine-grained human boundaries. Extensive experiments on TUM Dynamics and GTA-IM datasets demonstrate superior performance in dynamic human reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。