arXiv:2604.12929cs.CV2026-04被引 2

用高斯混合模型快速重建单目视频中动态手物交互,速度提升近40倍。

Grasp in Gaussians: Fast Monocular Reconstruction of Dynamic Hand-Object Interactions

  • 用轻量级高斯混合表示跟踪手与物体,结合预训练先验提高效率。
  • 在长序列上比以往方法快4.4至38.9倍,保持运动连贯性。
  • 适合需要实时交互重建的场景,如虚拟现实与机器人抓取模拟。

我们提出Grasp in Gaussians(GraG),一种从单目视频中快速稳健地重建动态3D手物交互的方法。不同于依赖复杂神经表示的近期方法,本方法在预训练大模型初始化后,高效追踪手与物体。核心思路是:利用强预训练的手与物体先验,通过经典跟踪文献中的紧凑高斯和(SoG)表示,恢复准确且时间稳定的运动。我们使用改进的SAM3D视频适配管道初始化物体姿态与几何,再通过子采样将密集高斯表示转为轻量级SoG。该紧凑表示支持高效快速追踪,同时保持几何保真度;预训练先验与简单几何/接触损失确保手物位置准确。对手部,采用互补策略:从现成的单目手姿态初始化出发,仅用2D关节、轮廓、深度与接触损失精修运动,避免逐帧优化详细3D手形,却维持稳定关节结构。在公开基准上的大量实验表明,GraG在长序列上重建手物交互的速度比之前工作快4.4至38.9倍,同时保持时间连贯性。

原文摘要 · Abstract (English)

We present Grasp in Gaussians (GraG), a fast and robust method for reconstructing dynamic 3D hand-object interactions from a single monocular video. Unlike recent approaches that optimize heavy neural representations, our method focuses on tracking the hand and the object efficiently, once initialized from pretrained large models. Our key insight is that, given strong pretrained object and hand priors, accurate and temporally stable hand-object motion can be recovered using a compact Sum-of-Gaussians (SoG) representation, revived from classical tracking literature and integrated with generative Gaussian-based initializations. We initialize object pose and geometry using a video-adapted SAM3D pipeline, then convert the resulting dense Gaussian representation into a lightweight SoG via subsampling. This compact representation enables efficient and fast tracking while preserving geometric fidelity, with the pretrained priors and simple geometric/contact losses providing accurate hand-object placement. For the hand, we adopt a complementary strategy: starting from off-the-shelf monocular hand pose initialization, we refine hand motion using simple yet effective 2D joint, silhouette, depth, and contact losses, avoiding per-frame refinement of a detailed 3D hand appearance model while maintaining stable articulation. Extensive experiments on public benchmarks demonstrate that GraG reconstructs temporally coherent hand-object interactions on long sequences 4.4x-38.9x faster than prior work while preserving temporally coherent hand-object motion.

动作捕捉手物交互高斯表示单目重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。