arXiv:2505.02405cs.ROcs.CV2025-05ICRA被引 3

基于信念场景图推断未见物体的空间分布,提升场景理解的常识性。

Estimating Commonsense Scene Composition on Belief Scene Graphs

  • 用联合概率建模物体在场景中的空间位置分布
  • 提出图神经网络与大模型结合的双版本框架
  • 适用于需要常识推理的室内场景理解任务

本文提出常识性场景构图概念,聚焦于通过估计未见物体的空间分布来扩展信念场景图。常识性场景构图指对场景中相关物体间空间关系的理解,本文将其建模为所有语义物体类别的可能位置的联合概率分布。所提框架包含两种相关性信息(CECI)模型变体:(i) 基于图卷积网络的基线方法;(ii) 融合基于大语言模型的空间本体的神经符号扩展。此外,文章详细描述了该任务的数据集生成过程。框架在模拟数据和真实室内环境中多次验证,展示了其在不同房间类型下进行空间场景解释的能力。

原文摘要 · Abstract (English)

This work establishes the concept of commonsense scene composition, with a focus on extending Belief Scene Graphs by estimating the spatial distribution of unseen objects. Specifically, the commonsense scene composition capability refers to the understanding of the spatial relationships among related objects in the scene, which in this article is modeled as a joint probability distribution for all possible locations of the semantic object class. The proposed framework includes two variants of a Correlation Information (CECI) model for learning probability distributions: (i) a baseline approach based on a Graph Convolutional Network, and (ii) a neuro-symbolic extension that integrates a spatial ontology based on Large Language Models (LLMs). Furthermore, this article provides a detailed description of the dataset generation process for such tasks. Finally, the framework has been validated through multiple runs on simulated data, as well as in a real-world indoor environment, demonstrating its ability to spatially interpret scenes across different room types.

场景理解空间推理大模型图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。