arXiv:2602.03200cs.CVcs.AI2026-02被引 1

首个单目视频实时重建手与场景4D动态关系的框架

Hand3R: Online 4D Hand-Scene Reconstruction in the Wild

  • 用视觉提示融合手部专家与场景基础模型
  • 单次前向传播实现高精度手形与场景几何重建
  • 适合需要实时交互理解的具身智能应用

对于具身人工智能而言,联合重建动态手部与密集场景上下文对理解物理交互至关重要。然而,现有方法大多在局部坐标系中恢复孤立的手部,忽略了周围三维环境。为此,我们提出Hand3R,首个从单目视频实现在线联合4D手-场景重建的框架。Hand3R通过场景感知的视觉提示机制,将预训练手部专家与4D场景基础模型协同整合。通过将高保真手部先验注入持久化场景记忆,该方法可在一次前向传播中同时重建精确的手部网格与稠密度量尺度的场景几何。实验表明,Hand3R摆脱了对离线优化的依赖,在局部手部重建与全局定位上均达到具有竞争力的性能。

原文摘要 · Abstract (English)

For Embodied AI, jointly reconstructing dynamic hands and the dense scene context is crucial for understanding physical interaction. However, most existing methods recover isolated hands in local coordinates, overlooking the surrounding 3D environment. To address this, we present Hand3R, the first online framework for joint 4D hand-scene reconstruction from monocular video. Hand3R synergizes a pre-trained hand expert with a 4D scene foundation model via a scene-aware visual prompting mechanism. By injecting high-fidelity hand priors into a persistent scene memory, our approach enables simultaneous reconstruction of accurate hand meshes and dense metric-scale scene geometry in a single forward pass. Experiments demonstrate that Hand3R bypasses the reliance on offline optimization and delivers competitive performance in both local hand reconstruction and global positioning.

4D重建手部建模具身智能单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。