arXiv:2603.23997cs.CV2026-03被引 2

无需标定相机,单张或多张图像都能精准重建3D手部网格。

HGGT: Robust and Flexible 3D Hand Mesh Reconstruction from Uncalibrated Images

  • 基于视觉-几何联合建模,直接从任意视角图像推断3D手形与相机位姿。
  • 在未标定的野外场景中表现优异,优于现有最先进方法。
  • 适合部署于普通摄像头,适用于机器人、VR/AR等实际应用。

从图像中恢复高保真3D手部几何是计算机视觉中的关键任务,在机器人、动画和虚拟现实等领域具有重要价值。可扩展的应用需兼顾精度与部署灵活性,要求能利用互联网上的海量非结构化图像数据,或在消费级RGB相机上无需复杂标定即可部署。然而,当前方法面临两难:单视图方法易于部署,但存在深度模糊与遮挡问题;多视图系统虽能解决不确定性,却通常依赖固定且标定好的采集环境,限制了真实场景的应用。为此,我们受3D基础模型启发,通过将任意视角下的手部重建重构为视觉-几何对齐任务,提出首个无需标定视角即可联合推理3D手部网格与相机位姿的前馈架构。大量实验表明,该方法在性能上超越现有基准,并展现出对未标定、野外场景的强大泛化能力。

原文摘要 · Abstract (English)

Recovering high-fidelity 3D hand geometry from images is a critical task in computer vision, holding significant value for domains such as robotics, animation and VR/AR. Crucially, scalable applications demand both accuracy and deployment flexibility, requiring the ability to leverage massive amounts of unstructured image data from the internet or enable deployment on consumer-grade RGB cameras without complex calibration. However, current methods face a dilemma. While single-view approaches are easy to deploy, they suffer from depth ambiguity and occlusion. Conversely, multi-view systems resolve these uncertainties but typically demand fixed, calibrated setups, limiting their real-world utility. To bridge this gap, we draw inspiration from 3D foundation models that learn explicit geometry directly from visual data. By reformulating hand reconstruction from arbitrary views as a visual-geometry grounded task, we propose a feed-forward architecture that, for the first time in literature, jointly infers 3D hand meshes and camera poses from uncalibrated views. Extensive evaluations show that our approach outperforms state-of-the-art benchmarks and demonstrates strong generalization to uncalibrated, in-the-wild scenarios. Here is the link of our project page: https://lym29.github.io/HGGT/.

3D重建手部建模无标定视觉几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。