单目视频重建松散衣物与手持物体交互的4D人体模型
DressRecon: Freeform 4D Human Reconstruction from Monocular Video
- 用神经隐式模型分离身体与衣物变形,结合通用人体先验与视频特定变形
- 在复杂衣物和物体交互场景下,重建质量优于现有方法
- 适合影视特效、虚拟试衣等需要高保真动态人体建模的场景
我们提出一种从单目视频中重建时序一致人体模型的方法,重点解决极松散衣物或手持物体交互的挑战。现有方法要么仅限于贴身衣物且无法处理物体交互,要么依赖校准的多视角捕获或个性化模板扫描,获取成本高。我们的核心思路是将大规模训练数据学习到的通用人体结构先验,与基于测试时优化拟合单个视频的“骨骼袋”式关节形变相结合。通过学习一个神经隐式模型,将身体与衣物变形作为独立运动层解耦。为捕捉衣物细微几何特征,优化过程中引入人体姿态、表面法线和光流等图像先验。最终生成的时间一致网格可进一步优化为显式3D高斯,支持高保真交互渲染。在具有高度挑战性的衣物变形与物体交互数据集上,DressRecon 的重建质量显著优于先前方法。
原文摘要 · Abstract (English)
We present a method to reconstruct time-consistent human body models from monocular videos, focusing on extremely loose clothing or handheld object interactions. Prior work in human reconstruction is either limited to tight clothing with no object interactions, or requires calibrated multi-view captures or personalized template scans which are costly to collect at scale. Our key insight for high-quality yet flexible reconstruction is the careful combination of generic human priors about articulated body shape (learned from large-scale training data) with video-specific articulated "bag-of-bones" deformation (fit to a single video via test-time optimization). We accomplish this by learning a neural implicit model that disentangles body versus clothing deformations as separate motion model layers. To capture subtle geometry of clothing, we leverage image-based priors such as human body pose, surface normals, and optical flow during optimization. The resulting neural fields can be extracted into time-consistent meshes, or further optimized as explicit 3D Gaussians for high-fidelity interactive rendering. On datasets with highly challenging clothing deformations and object interactions, DressRecon yields higher-fidelity 3D reconstructions than prior art. Project page: https://jefftan969.github.io/dressrecon/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。