arXiv:2502.14004cs.GRcs.LG2025-02IJCAI被引 3

构建交互式3D物体重建新基准,解决多部件状态建模难题

Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object Reconstruction

  • 用空间差异张量高效建模可动部件的所有离散状态
  • 在仅观察单个部件状态下实现未见组合状态的还原
  • 适合研究人机交互3D重建的学者与工程师

近期隐式3D重建方法(如神经渲染场和高斯溅射)主要聚焦于静态或连续运动物体的新视角合成,但难以高效建模具有n个可动部件的人机交互物体,需2^n个独立模型表示所有离散状态。为此,我们提出Inter3D,一个针对人机交互物体的新基准与方法。构建自采数据集,涵盖常见交互物体,并设计新评估流程:训练时仅观察单个部件状态,组合状态完全未见。提出强基线方法,利用空间差异张量高效建模所有状态;为缓解训练中相机轨迹约束过强问题,引入互状态正则化机制以提升可动部件的空间密度一致性;探索两种体素网格采样策略以提高训练效率。在所提基准上开展大量实验,揭示任务挑战并验证方法优越性。

原文摘要 · Abstract (English)

Recent advancements in implicit 3D reconstruction methods, e.g., neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these approaches struggle to efficiently model a human-interactive object with n movable parts, requiring 2^n separate models to represent all discrete states. To overcome this limitation, we propose Inter3D, a new benchmark and approach for novel state synthesis of human-interactive objects. We introduce a self-collected dataset featuring commonly encountered interactive objects and a new evaluation pipeline, where only individual part states are observed during training, while part combination states remain unseen. We also propose a strong baseline approach that leverages Space Discrepancy Tensors to efficiently modelling all states of an object. To alleviate the impractical constraints on camera trajectories across training states, we propose a Mutual State Regularization mechanism to enhance the spatial density consistency of movable parts. In addition, we explore two occupancy grid sampling strategies to facilitate training efficiency. We conduct extensive experiments on the proposed benchmark, showcasing the challenges of the task and the superiority of our approach.

3D重建人机交互隐式表示基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。