利用几何先验提升单目手部三维重建精度
GeoHand: Unlocking Prior Geometry Knowledge for Monocular 3D Hand Reconstruction

- 从冻结的几何估计器提取高质量先验并适配手部细节
- 在三数据集上实现当前最优,尤其在遮挡和交互场景下
- 适合需要高精度手部重建的研究者与工业应用
单目3D手部重建本质上是几何问题,但仅依赖RGB外观特征难以解决自遮挡和手物交互带来的严重歧义。虽然深度信息可提供空间线索,但原始传感器捕获的深度图噪声大且不完整,限制了其在精细重建中的应用。为此,我们提出GeoHand,一种从冻结的通用单目几何估计器(MoGe2)中解锁高质量几何先验的新框架。由于这些先验面向通用场景,我们引入地图级GeoAdapter重新校准空间特征,使其专用于手部重建。为进一步系统融合这些适配后的先验而不压倒固有的RGB外观线索,采用门控跨模态令牌融合策略。最后,为确保精确的局部关节结构,设计了基于关键点查询的迭代精修器(KQIR),利用投影关节位置查询几何感知图像特征进行空间修正。通过统一管道结合全局几何消歧与局部精修,GeoHand在FreiHAND、DexYCB和HO3Dv3数据集上达到当前最优性能,尤其在严重遮挡和手物交互情况下表现突出。
原文摘要 · Abstract (English)
Monocular 3D hand reconstruction is intrinsically a geometric problem, yet RGB appearance features alone often struggle to resolve severe ambiguities caused by self-occlusions and hand-object interactions. While introducing depth can explicitly provide spatial cues, raw sensor-captured depth maps are extensively noisy and incomplete, limiting their usefulness for fine-grained hand reconstruction. To bridge this gap, we propose GeoHand, a novel framework that unlocks high-quality geometric priors from a frozen foundational monocular geometry estimator (MoGe2). Recognizing that these priors are oriented toward general scenes, we introduce a map-level GeoAdapter to recalibrate the spatial features, specifically adapting them for detailed hand reconstruction. Furthermore, to systematically integrate these adapted priors without overwhelming intrinsic RGB appearance cues, we employ a gated cross-modal token fusion strategy. Finally, to secure precise local articulation, we design a Keypoint-Queried Iterative Refiner (KQIR) that uses projected joint locations to query geometry-aware image features for spatial correction. By combining global geometric disambiguation with local refinement in a unified pipeline, GeoHand achieves state-of-the-art performance on FreiHAND, DexYCB, and HO3Dv3, especially under severe occlusions and hand-object interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。