用单视角模型做多视角人体建模,无需标定也能高精度重建。
Monocular Models are Strong Learners for Multi-View Human Mesh Recovery
- 利用预训练单视角模型作为先验,无需多视角训练数据。
- 通过测试时优化提升重建一致性与解剖合理性,性能超越有监督多视角模型。
- 适合缺乏多视角数据或需快速部署的场景,如实时捕捉与移动设备应用。
多视角人体网格恢复(HMR)在多个领域广泛应用,对高精度和强泛化能力要求严格。现有方法分为基于几何和基于学习两类:前者依赖繁琐的相机标定,后者因缺少多视角训练数据,泛化能力差,难以适应真实场景。为实现免标定、任意相机布局下的重建,我们提出一种无需训练的框架,利用预训练的单视角HMR模型作为强先验,避免多视角训练数据需求。该方法首先从单视角预测构建鲁棒且一致的多视角初始估计,再通过测试时优化,在多视角一致性与解剖约束下进行精细化调整。大量实验表明,该方法在标准基准上达到领先性能,优于显式多视角监督训练的模型。
原文摘要 · Abstract (English)
Multi-view human mesh recovery (HMR) is broadly deployed in diverse domains where high accuracy and strong generalization are essential. Existing approaches can be broadly grouped into geometry-based and learning-based methods. However, geometry-based methods (e.g., triangulation) rely on cumbersome camera calibration, while learning-based approaches often generalize poorly to unseen camera configurations due to the lack of multi-view training data, limiting their performance in real-world scenarios. To enable calibration-free reconstruction that generalizes to arbitrary camera setups, we propose a training-free framework that leverages pretrained single-view HMR models as strong priors, eliminating the need for multi-view training data. Our method first constructs a robust and consistent multi-view initialization from single-view predictions, and then refines it via test-time optimization guided by multi-view consistency and anatomical constraints. Extensive experiments demonstrate state-of-the-art performance on standard benchmarks, surpassing multi-view models trained with explicit multi-view supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。