无需模板标记,通过多视角视频实现高保真4D手物交互重建。
High-Fidelity 4D Hand-Object Capture via Multi-View Spatiotemporal Tracking and Physics-Aware Gaussians

- 用多视角时空变换器融合几何与时间信息,提供精准初始姿态和物体结构。
- 基于物理约束的高斯优化框架,消除碰撞并提升视觉真实性。
- 适合需要自动化4D内容生成的研究者和开发者使用。
在具身智能与空间计算中,对高保真4D手物交互(HOI)数据的需求日益增长,但当前方法依赖预扫描物体模板和物理标记,成为瓶颈。尽管近期方法已能从视频中重建4D HOI,但仍高度依赖手部与物体姿态的初始估计。而图像中姿态估计在严重遮挡下尤为困难,而这正是手物交互的固有特性。本文提出一种新系统,仅需同步校准的多视角视频,无需任何模板或标记,即可实现鲁棒且精确的双手与物体重建。系统包含两大创新组件:(1) 多视角前馈变换器模型,通过融合跨视角几何与时间线索,为手部与物体提供一致度量的可靠初始姿态及密集物体几何;(2) 手物物理感知的高斯优化框架,集成四面体约束、碰撞精修与外观分解,生成物理合理且视觉准确的重建结果。在公开基准和大规模内部数据集上验证,本方法实现无伪影、鲁棒性强的重建,为自动化4D资产生成提供高效基础。
原文摘要 · Abstract (English)
The growing demand for high-fidelity 4D hand-object interaction (HOI) data in embodied AI and spatial computing is currently bottlenecked by the reliance on pre-scanned object templates and physical markers. While recent methods have demonstrated promising results in reconstructing 4D hand-object interaction from videos, they are highly sensitive to initial estimates of hand and object poses. Yet, estimating these poses from images is challenging, in particular under severe occlusion which is inherent in hand-object interaction scenarios. We propose a novel system for the robust and accurate reconstruction of hands and objects from synchronized and calibrated multi-view videos without requiring any templates or markers. Our system consists of two main components with key innovations: (1) a multi-view feed-forward transformer model that aggregates cross-view geometry and temporal cues to provide a reliable, metric-consistent initialization for both poses and dense object geometry, and (2) a hand-object physics-aware Gaussian-based optimization framework to refine the initial estimates, integrating tetrahedral constraints, collision refinement, and appearance decomposition to produce physically plausible and visually accurate reconstruction. Validated on public benchmarks and an extensive internal dataset, our pipeline achieves highly robust, artifact-free reconstruction, providing an efficient foundation for automated 4D asset generation. Our project page are available at https://zyshen021.github.io/HOSTPG/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。