融合视觉与符号推理,提升机器人在复杂环境中的空间理解能力
A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
- 结合全景图与3D点云,用神经网络感知+符号逻辑建模空间关系
- 在JRDB-Reasoning数据集上表现优于现有模型,尤其在密集人机环境中
- 结构化场景图支持可解释查询,适合机器人与具身智能应用
空间推理是机器人领域中理解物体间关系与交互的高阶认知任务。现有视觉语言模型虽擅长感知,但因依赖图像相关性、缺乏显式逻辑推理,在精细空间理解上表现不足。本文提出一种新型神经符号框架,融合全景图像与3D点云信息,结合神经感知与符号推理,显式建模空间与逻辑关系。框架包含感知模块(检测实体并提取属性)和推理模块(构建结构化场景图以支持精确可解释查询)。在JRDB-Reasoning数据集上的评估显示,该方法在人群密集的人造环境中表现更优,且模型轻量,适用于机器人与具身智能应用。
原文摘要 · Abstract (English)
Visual reasoning, particularly spatial reasoning, is a challenging cognitive task that requires understanding object relationships and their interactions within complex environments, especially in robotics domain. Existing vision_language models (VLMs) excel at perception tasks but struggle with fine-grained spatial reasoning due to their implicit, correlation-driven reasoning and reliance solely on images. We propose a novel neuro_symbolic framework that integrates both panoramic-image and 3D point cloud information, combining neural perception with symbolic reasoning to explicitly model spatial and logical relationships. Our framework consists of a perception module for detecting entities and extracting attributes, and a reasoning module that constructs a structured scene graph to support precise, interpretable queries. Evaluated on the JRDB-Reasoning dataset, our approach demonstrates superior performance and reliability in crowded, human_built environments while maintaining a lightweight design suitable for robotics and embodied AI applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。