用图像到视频扩散模型实现人形三维几何的精准时序一致估计
GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion
- 先用图像模型估首帧深度法线,再用视频扩散模型生成后续帧
- 在仅需少量4D数据下,实现时序一致且细节丰富的3D重建
- 适合需要高保真人体动态建模的研究者与应用开发者
从单目视频中准确估计三维人体几何并保持时序一致性是计算机视觉中的难题。现有方法主要针对单图优化,常导致时序不一致且难以捕捉细微动态。为解决这一问题,我们提出GeoMan,一种新架构,可从单目人体视频中生成精确且时序一致的深度与法线图。针对高质量4D训练数据稀缺的问题,GeoMan采用图像模型先估计视频首帧的深度与法线,再以之为条件引导视频扩散模型,将几何估计任务转化为图像到视频生成问题。该设计将几何推断负担交给图像模型,视频模型则专注于细节生成,并利用大规模视频数据学习的先验知识。此外,为实现准确的人体尺度建模,我们引入根相对深度表示法,保留关键人体尺度信息,更易从单目输入中估计,克服传统仿射不变与度量深度表示的局限。GeoMan在定性与定量评估中均达到当前最佳性能,有效解决了人体视频三维几何估计的长期挑战。
原文摘要 · Abstract (English)
Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies and fail to capture fine-grained dynamic details. To address these limitations, we present GeoMan, a novel architecture designed to produce accurate and temporally consistent depth and normal estimations from monocular human videos. GeoMan addresses two key challenges: the scarcity of high-quality 4D training data and the need for metric depth estimation to accurately model human size. To overcome the first challenge, GeoMan employs an image-based model to estimate depth and normals for the first frame of a video, which then conditions a video diffusion model, reframing video geometry estimation task as an image-to-video generation problem. This design offloads the heavy lifting of geometric estimation to the image model and simplifies the video model's role to focus on intricate details while using priors learned from large-scale video datasets. Consequently, GeoMan improves temporal consistency and generalizability while requiring minimal 4D training data. To address the challenge of accurate human size estimation, we introduce a root-relative depth representation that retains critical human-scale details and is easier to be estimated from monocular inputs, overcoming the limitations of traditional affine-invariant and metric depth representations. GeoMan achieves state-of-the-art performance in both qualitative and quantitative evaluations, demonstrating its effectiveness in overcoming longstanding challenges in 3D human geometry estimation from videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。