arXiv:2507.08416cs.CV2025-07ICCV被引 10

让机器人像人一样识别并补全被遮挡的3D物体

InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes

  • 通过跨视角空间对比学习增强遮挡场景下的语义监督
  • 在真实与合成场景中实现高精度物体分解与完整重建
  • 适合需要精准3D感知的机器人、自动驾驶领域

人类能自然识别并心理补全杂乱环境中的遮挡物体,但赋予机器人类似认知能力仍具挑战。现有重建技术将场景视为无差别的整体,难以从局部观测中识别完整物体。本文提出InstaScene,一种面向复杂场景的3D实例分解与重建新范式,目标是同时实现任意实例的精确分解与完整重建。为提升分解精度,我们设计了一种新颖的跨视角光栅化追踪空间对比学习,显著增强复杂场景中的语义监督。为克服观测不完整问题,引入原位生成机制,利用有效观测与几何线索,引导3D生成模型重建与真实世界无缝契合的完整物体。在真实与合成复杂场景上的实验表明,该方法在实例分解准确率上表现优异,且重建物体具有高几何保真度与视觉完整性。

原文摘要 · Abstract (English)

Humans can naturally identify and mentally complete occluded objects in cluttered environments. However, imparting similar cognitive ability to robotics remains challenging even with advanced reconstruction techniques, which models scenes as undifferentiated wholes and fails to recognize complete object from partial observations. In this paper, we propose InstaScene, a new paradigm towards holistic 3D perception of complex scenes with a primary goal: decomposing arbitrary instances while ensuring complete reconstruction. To achieve precise decomposition, we develop a novel spatial contrastive learning by tracing rasterization of each instance across views, significantly enhancing semantic supervision in cluttered scenes. To overcome incompleteness from limited observations, we introduce in-situ generation that harnesses valuable observations and geometric cues, effectively guiding 3D generative models to reconstruct complete instances that seamlessly align with the real world. Experiments on scene decomposition and object completion across complex real-world and synthetic scenes demonstrate that our method achieves superior decomposition accuracy while producing geometrically faithful and visually intact objects.

3D感知实例分割生成建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。