无需预设物体类别,单目视频实现人体与物体4D动态重建
CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction
- 基于基础模型预测,通过可学习的渲染对比机制联合优化姿态
- 在分布内数据上误差降低38%,分布外数据上降低36%
- 零样本泛化至真实网络视频,适用于未知物体类别场景
从常见传感器如RGB相机准确捕捉人体-物体交互对人类理解、游戏和机器人学习等应用至关重要。然而,仅凭单视角RGB图像推断4D交互极具挑战,受限于未知物体与人体信息、深度模糊、遮挡及复杂运动,导致三维与时间维度重建不一致。现有方法常依赖真值物体模板或限定物体类别。本文提出CARI4D,首个无类别依赖的方法,可从单目RGB视频中以度量尺度重建时空一致的4D人体-物体交互。我们设计了一种姿态假设选择算法,鲁棒融合基础模型个体预测,并通过可学习的渲染-比对范式联合优化,确保空间、时间与像素对齐;最终通过物理约束推理细化复杂接触关系。实验表明,本方法在分布内数据上重建误差降低38%,分布外数据上降低36%。模型具备超越训练类别的泛化能力,可零样本应用于真实互联网视频。代码与预训练模型将公开发布。
原文摘要 · Abstract (English)
Accurate capture of human-object interaction from ubiquitous sensors like RGB cameras is important for applications in human understanding, gaming, and robot learning. However, inferring 4D interactions from a single RGB view is highly challenging due to the unknown object and human information, depth ambiguity, occlusion, and complex motion, which hinder consistent 3D and temporal reconstruction. Previous methods simplify the setup by assuming ground truth object template or constraining to a limited set of object categories. We present CARI4D, the first category-agnostic method that reconstructs spatially and temporarily consistent 4D human-object interaction at metric scale from monocular RGB videos. To this end, we propose a pose hypothesis selection algorithm that robustly integrates the individual predictions from foundation models, jointly refine them through a learned render-and-compare paradigm to ensure spatial, temporal and pixel alignment, and finally reasoning about intricate contacts for further refinement satisfying physical constraints. Experiments show that our method outperforms prior art by 38% on in-distribution dataset and 36% on unseen dataset in terms of reconstruction error. Our model generalizes beyond the training categories and thus can be applied zero-shot to in-the-wild internet videos. Our code and pretrained models will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。