用深度信息提升单目视频人体建模的精度与稳定性。
Depth-Guided Metric-Aware Temporal Consistency for Monocular Video Human Mesh Recovery
- 融合深度与图像特征,自适应整合几何先验
- 在三个基准上显著降低遮挡影响,提升空间精度
- 适合需要高精度人体动态重建的研究者
单目视频人体网格恢复因深度模糊和尺度不确定面临度量一致性与时间稳定性的根本挑战。现有方法依赖RGB特征和时间平滑,难以应对深度排序错误、尺度漂移及遮挡引发的不稳定性。本文提出一种深度引导的综合框架,通过三个协同组件实现度量感知的时间一致性:深度引导的多尺度融合模块,通过置信度感知门控自适应融合几何先验与RGB特征;深度校准的度量感知姿态与形状(D-MAPS)估计器,利用深度校准的骨骼统计量实现尺度一致的初始化;运动-深度对齐优化(MoDAR)模块,通过运动动态与几何线索间的跨模态注意力强制时间一致性。该方法在三个挑战性基准上取得优异表现,显著提升对严重遮挡的鲁棒性与空间精度,同时保持计算效率。
原文摘要 · Abstract (English)
Monocular video human mesh recovery faces fundamental challenges in maintaining metric consistency and temporal stability due to inherent depth ambiguities and scale uncertainties. While existing methods rely primarily on RGB features and temporal smoothing, they struggle with depth ordering, scale drift, and occlusion-induced instabilities. We propose a comprehensive depth-guided framework that achieves metric-aware temporal consistency through three synergistic components: A Depth-Guided Multi-Scale Fusion module that adaptively integrates geometric priors with RGB features via confidence-aware gating; A Depth-guided Metric-Aware Pose and Shape (D-MAPS) estimator that leverages depth-calibrated bone statistics for scale-consistent initialization; A Motion-Depth Aligned Refinement (MoDAR) module that enforces temporal coherence through cross-modal attention between motion dynamics and geometric cues. Our method achieves superior results on three challenging benchmarks, demonstrating significant improvements in robustness against heavy occlusion and spatial accuracy while maintaining computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。