让3D场景识别重复物品,提升重建与分割精度
Lookalike3D: Seeing Double in 3D
- 用多视角图像和大模型语义先验区分物体是否相似
- 在76,000对物体上实现比基线高104%的IoU
- 适合做3D重建、部件共分割等任务的研究者
3D物体理解与生成方法已取得显著成果,但常忽略真实场景中广泛存在的重复物体信息。本文提出室内场景中看似物体检测任务,利用完全相同或近似相同的物体对之间的重复与互补线索。给定一个输入场景,任务是基于多视角图像判断物体对为相同、相似或不同。为此,我们提出Lookalike3D,一种多视角图像变换器,通过利用大型图像基础模型的强大语义先验,有效区分此类物体对。为支持该任务,我们构建了3DTwins数据集,包含基于ScanNet++的76,000个手动标注的相同、相似和不同物体对,并在该数据集上实现比基线高出104%的IoU。我们还展示了该方法如何提升下游任务,如联合3D物体重建与部件共分割,将重复物体转化为高质量3D感知的强大线索。代码、数据集和模型将公开发布。
原文摘要 · Abstract (English)
3D object understanding and generation methods produce impressive results, yet they often overlook a pervasive source of information in real-world scenes: repeated objects. We introduce the task of lookalike object detection in indoor scenes, which leverages repeated and complementary cues from identical and near-identical object pairs. Given an input scene, the task is to classify pairs of objects as identical, similar or different using multiview images as input. To address this, we present Lookalike3D, a multiview image transformer that effectively distinguishes such object pairs by harnessing strong semantic priors from large image foundation models. To support this task, we collected the 3DTwins dataset, containing 76k manually annotated identical, similar and different pairs of objects based on ScanNet++, and show an improvement of 104% IoU over baselines. We demonstrate how our method improves downstream tasks such as enabling joint 3D object reconstruction and part co-segmentation, turning repeated and lookalike objects into a powerful cue for consistent, high-quality 3D perception. Our code, dataset and models will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。