arXiv:2603.16868cs.CVcs.AI2026-03

构建真实杂乱厨房数据集,实现高精度物体间物理接触的3D场景重建。

MessyKitchens: Contact-rich object-level 3D scene reconstruction

  • 基于SAM 3D扩展多物体解码器,联合重建多个物体的3D形状与位置。
  • 在新数据集上物体对齐误差降低47%,物体穿透率下降62%。
  • 适合机器人抓取、动画制作等需物理合理交互的场景重建任务。

单目3D场景重建近年来取得显著进展,得益于现代神经架构和大规模数据,深度估计性能大幅提升。然而,将常见场景分解为独立3D物体仍面临巨大挑战,主要源于物体种类繁多、频繁遮挡及复杂关系。尤其在机器人和动画应用中,还需满足非穿透性与真实接触的物理合理性。本文从两个方向推进物体级场景重建:首先,提出新数据集MessyKitchens,包含真实杂乱厨房场景,提供高保真物体级真值,涵盖3D形状、位姿及精确物体接触信息;其次,在SAM 3D单物体重建基础上,引入多物体解码器(MOD),实现联合物体级场景重建。通过对比验证,该数据集在注册精度与物体间穿透率方面显著优于现有数据集;在三个数据集上,MOD方法均一致且显著超越当前最优水平。项目代码与预训练模型将公开于https://messykitchens.github.io/。

原文摘要 · Abstract (English)

Monocular 3D scene reconstruction has recently seen significant progress. Powered by the modern neural architectures and large-scale data, recent methods achieve high performance in depth estimation from a single image. Meanwhile, reconstructing and decomposing common scenes into individual 3D objects remains a hard challenge due to the large variety of objects, frequent occlusions and complex object relations. Notably, beyond shape and pose estimation of individual objects, applications in robotics and animation require physically-plausible scene reconstruction where objects obey physical principles of non-penetration and realistic contacts. In this work we advance object-level scene reconstruction along two directions. First, we introduceMessyKitchens, a new dataset with real-world scenes featuring cluttered environments and providing high-fidelity object-level ground truth in terms of 3D object shapes, poses and accurate object contacts. Second, we build on the recent SAM 3D approach for single-object reconstruction and extend it with Multi-Object Decoder (MOD) for joint object-level scene reconstruction. To validate our contributions, we demonstrate MessyKitchens to significantly improve previous datasets in registration accuracy and inter-object penetration. We also compare our multi-object reconstruction approach on three datasets and demonstrate consistent and significant improvements of MOD over the state of the art. Our new benchmark, code and pre-trained models will become publicly available on our project website: https://messykitchens.github.io/.

3D重建物体接触机器人数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。