arXiv:2410.05869cs.CVcs.AI2024-10CVPR被引 8

用生成模型检测图像外或被遮挡的物体,让看不见的也能被发现。

Believing is Seeing: Unobserved Object Detection using Generative Models

  • 利用预训练生成模型推断被遮挡或在画面外的物体位置。
  • 在RealEstate10k和NYU Depth v2数据集上实现有效检测,验证可行性。
  • 适合做机器人感知、自动驾驶等需全场景理解的任务。

能否检测图像中不可见但位于相机附近的物体?本文提出2D、2.5D和3D未观测物体检测新任务,旨在预测被遮挡或位于图像边界外的物体位置。我们适配了多种先进的预训练生成模型,包括2D与3D扩散模型及视觉-语言模型,并证明它们可有效推断未直接观测到的物体存在。为评估该任务,我们设计了一套涵盖多维度性能的评测指标。在RealEstate10k与NYU Depth v2数据集上的室内场景实证表明,生成模型在未观测物体检测任务中表现良好,具备应用潜力。

原文摘要 · Abstract (English)

Can objects that are not visible in an image -- but are in the vicinity of the camera -- be detected? This study introduces the novel tasks of 2D, 2.5D and 3D unobserved object detection for predicting the location of nearby objects that are occluded or lie outside the image frame. We adapt several state-of-the-art pre-trained generative models to address this task, including 2D and 3D diffusion models and vision-language models, and show that they can be used to infer the presence of objects that are not directly observed. To benchmark this task, we propose a suite of metrics that capture different aspects of performance. Our empirical evaluation on indoor scenes from the RealEstate10k and NYU Depth v2 datasets demonstrate results that motivate the use of generative models for the unobserved object detection task.

生成模型目标检测未观测物体视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。