无需关键点检测,直接从单目视频重建手物3D姿态与形状
HOSt3R: Keypoint-free Hand-Object 3D Reconstruction from RGB images

- 提出无关键点的单目视频手物3D变换估计方法
- 在SHOWMe上达到当前最佳性能,误差低于15.6mm
- 适用于未见物体类别,无需标定或模板
手物3D重建在人机交互和沉浸式AR/VR中日益重要。现有方法通常采用两阶段流程:先进行手物3D跟踪,再进行多视角重建。但这些方法依赖于关键点检测技术(如SfM和手部关键点优化),在物体几何多样、纹理弱或手物相互遮挡时表现不佳,限制了泛化能力。为此,本文提出一种无需关键点检测的鲁棒方法,直接从单目运动视频中估计手物3D变换,并结合多视角重建流程准确恢复手物3D形状。所提方法HOSt3R无需预扫描物体模板或相机内参,在SHOWMe基准上实现当前最优性能,3D位置误差低于15.6mm;同时在HO3D数据集序列上验证了对未见物体类别的良好泛化能力。
原文摘要 · Abstract (English)
Hand-object 3D reconstruction has become increasingly important for applications in human-robot interaction and immersive AR/VR experiences. A common approach for object-agnostic hand-object reconstruction from RGB sequences involves a two-stage pipeline: hand-object 3D tracking followed by multi-view 3D reconstruction. However, existing methods rely on keypoint detection techniques, such as Structure from Motion (SfM) and hand-keypoint optimization, which struggle with diverse object geometries, weak textures, and mutual hand-object occlusions, limiting scalability and generalization. As a key enabler to generic and seamless, non-intrusive applicability, we propose in this work a robust, keypoint detector-free approach to estimating hand-object 3D transformations from monocular motion video/images. We further integrate this with a multi-view reconstruction pipeline to accurately recover hand-object 3D shape. Our method, named HOSt3R, is unconstrained, does not rely on pre-scanned object templates or camera intrinsics, and reaches state-of-the-art performance for the tasks of object-agnostic hand-object 3D transformation and shape estimation on the SHOWMe benchmark. We also experiment on sequences from the HO3D dataset, demonstrating generalization to unseen object categories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。