arXiv:2602.04053cs.CV2026-02被引 1

通过逐次移除物体,从单张图像重建复杂场景的3D结构。

Seeing Through Clutter: Structured 3D Scene Reconstruction via Iterative Object Removal

  • 用大模型逐个检测并移除前景物体,分步简化场景
  • 在3D-Front和ADE20K上达到当前最优鲁棒性
  • 无需特定训练,直接受益于大模型进展

我们提出SeeingThroughClutter,一种从单张图像重构结构化3D场景的方法,通过逐个分割和建模物体实现。以往方法依赖语义分割和深度估计等中间任务,在复杂场景尤其是遮挡和杂乱情况下表现不佳。为此,我们设计了一种迭代物体移除与重建流程,将复杂场景分解为一系列更简单的子任务。利用视觉语言模型(VLMs)作为调度器,依次执行检测、分割、物体移除和3D拟合。实验表明,移除物体后能获得更清晰的后续物体分割结果,即使在高度遮挡场景中也有效。该方法无需任务特定训练,可直接利用基础模型的持续进步。我们在3D-Front和ADE20K数据集上展示了最先进的鲁棒性。

原文摘要 · Abstract (English)

We present SeeingThroughClutter, a method for reconstructing structured 3D representations from single images by segmenting and modeling objects individually. Prior approaches rely on intermediate tasks such as semantic segmentation and depth estimation, which often underperform in complex scenes, particularly in the presence of occlusion and clutter. We address this by introducing an iterative object removal and reconstruction pipeline that decomposes complex scenes into a sequence of simpler subtasks. Using VLMs as orchestrators, foreground objects are removed one at a time via detection, segmentation, object removal, and 3D fitting. We show that removing objects allows for cleaner segmentations of subsequent objects, even in highly occluded scenes. Our method requires no task-specific training and benefits directly from ongoing advances in foundation models. We demonstrate stateof-the-art robustness on 3D-Front and ADE20K datasets. Project Page: https://rioak.github.io/seeingthroughclutter/

3D重建视觉语言模型物体移除单图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。