利用人体外观时间一致性,实现每秒12帧的快速3D建模。
Link to the Past: Temporal Propagation for Fast 3D Human Reconstruction from Monocular Video
- 通过维护统一外观表示,减少视频流中的重复计算。
- 在标准数据集上达到或超越现有方法质量,最高达12帧/秒。
- 适合需要实时处理的3D人体重建场景。
从单目视频中快速重建穿着衣物的人体3D模型仍是计算机视觉中的重大挑战,尤其在于计算效率与重建质量之间的平衡。当前方法要么专注于静态图像重建但计算量过大,要么通过逐视频优化实现高质量,但需数分钟至数小时处理时间,无法满足实时应用需求。为此,我们提出TemPoFast3D,一种利用人体外观时间一致性的新方法,在保持重建质量的同时减少冗余计算。该方法为即插即用式设计,通过高效的坐标映射机制,将像素对齐的重建网络改造为可处理连续视频流,持续维护并优化一个基准外观表示。大量实验表明,TemPoFast3D在标准指标上匹配或超越现有最佳方法,可在多种姿态与外观下实现高质量纹理重建,最大速度达12帧/秒。
原文摘要 · Abstract (English)
Fast 3D clothed human reconstruction from monocular video remains a significant challenge in computer vision, particularly in balancing computational efficiency with reconstruction quality. Current approaches are either focused on static image reconstruction but too computationally intensive, or achieve high quality through per-video optimization that requires minutes to hours of processing, making them unsuitable for real-time applications. To this end, we present TemPoFast3D, a novel method that leverages temporal coherency of human appearance to reduce redundant computation while maintaining reconstruction quality. Our approach is a "plug-and play" solution that uniquely transforms pixel-aligned reconstruction networks to handle continuous video streams by maintaining and refining a canonical appearance representation through efficient coordinate mapping. Extensive experiments demonstrate that TemPoFast3D matches or exceeds state-of-the-art methods across standard metrics while providing high-quality textured reconstruction across diverse pose and appearance, with a maximum speed of 12 FPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。