从单目视频重建双手与未知物体的3D交互,突破遮挡难题。
BIGS: Bimanual Category-agnostic Interaction Reconstruction from Monocular Videos via 3D Gaussian Splatting

- 用3D高斯点云结合扩散模型先验重建未知物体
- 共享高斯表示双手,解决视角受限问题,提升精度
- 新增交互优化步骤,实现手物3D对齐,适合动作分析场景
手-物交互(HOI)三维重建是基础性任务,应用广泛。现有方法难以处理单目视频中双手与未知物体交互的复杂情况,尤其受双重手部遮挡影响。本文提出BIGS方法,通过3D高斯点云重建双手及未知物体。为克服物体部分被遮挡的问题,利用预训练扩散模型结合得分蒸馏采样(SDS)损失,恢复未见物体区域。对于双手,采用人体手部模型(MANO)的3D先验,并共享单一高斯表示以累积有限视角下的手部信息。进一步在高斯优化过程中引入交互主体对齐优化步骤,增强手物间3D空间一致性。在两个挑战性数据集上,本方法在3D手姿态估计(MPJPE)、3D物体重建(CDh, CDo, F10)和渲染质量(PSNR, SSIM, LPIPS)指标上均达到当前最优水平。
原文摘要 · Abstract (English)
Reconstructing 3Ds of hand-object interaction (HOI) is a fundamental problem that can find numerous applications. Despite recent advances, there is no comprehensive pipeline yet for bimanual class-agnostic interaction reconstruction from a monocular RGB video, where two hands and an unknown object are interacting with each other. Previous works tackled the limited hand-object interaction case, where object templates are pre-known or only one hand is involved in the interaction. The bimanual interaction reconstruction exhibits severe occlusions introduced by complex interactions between two hands and an object. To solve this, we first introduce BIGS (Bimanual Interaction 3D Gaussian Splatting), a method that reconstructs 3D Gaussians of hands and an unknown object from a monocular video. To robustly obtain object Gaussians avoiding severe occlusions, we leverage prior knowledge of pre-trained diffusion model with score distillation sampling (SDS) loss, to reconstruct unseen object parts. For hand Gaussians, we exploit the 3D priors of hand model (i.e., MANO) and share a single Gaussian for two hands to effectively accumulate hand 3D information, given limited views. To further consider the 3D alignment between hands and objects, we include the interacting-subjects optimization step during Gaussian optimization. Our method achieves the state-of-the-art accuracy on two challenging datasets, in terms of 3D hand pose estimation (MPJPE), 3D object reconstruction (CDh, CDo, F10), and rendering quality (PSNR, SSIM, LPIPS), respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。