arXiv:2410.07394cs.CV2024-10被引 4

用3D几何特征提升机器人对物体空间关系的理解能力

Structured Spatial Reasoning with Open Vocabulary Object Detectors

  • 融合3D几何信息与开放词汇检测器进行结构化概率推理
  • 在真实和合成数据集上比主流VLMs高出20%以上
  • 适合需要精准空间感知的机器人任务开发者

空间关系推理对机器人执行抓取、重排和搜寻等任务至关重要。准确识别物体及其位置是完成这些任务的关键。近期工作利用先进的视觉语言模型(VLMs)提升了这一能力。本文提出一种结构化的概率方法,将丰富的3D几何特征与最先进的开放词汇物体检测器结合,增强机器人感知中的空间推理能力。我们在真实世界的RGB-D主动视觉数据集[1]中标注了空间短语,并在该数据集及合成的语义抽象数据集[2]上进行了实验。结果表明,所提方法在空间关系定位上显著优于当前开源VLMs,性能提升超过20%。

原文摘要 · Abstract (English)

Reasoning about spatial relationships between objects is essential for many real-world robotic tasks, such as fetch-and-delivery, object rearrangement, and object search. The ability to detect and disambiguate different objects and identify their location is key to successful completion of these tasks. Several recent works have used powerful Vision and Language Models (VLMs) to unlock this capability in robotic agents. In this paper we introduce a structured probabilistic approach that integrates rich 3D geometric features with state-of-the-art open-vocabulary object detectors to enhance spatial reasoning for robotic perception. The approach is evaluated and compared against zero-shot performance of the state-of-the-art Vision and Language Models (VLMs) on spatial reasoning tasks. To enable this comparison, we annotate spatial clauses in real-world RGB-D Active Vision Dataset [1] and conduct experiments on this and the synthetic Semantic Abstraction [2] dataset. Results demonstrate the effectiveness of the proposed method, showing superior performance of grounding spatial relations over state of the art open-source VLMs by more than 20%.

空间推理机器人感知开放词汇检测3D几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。