arXiv:2504.18380cs.SEcs.AI2025-04被引 1

将3D几何信息转化为可查询的语义知识,提升XR场景理解能力。

Spatial Reasoner: A 3D Inference Pipeline for XR Applications

  • 基于定向3D框与形式化空间谓词构建推理框架
  • 支持'在...上''在...后'等10余类空间关系动态判断
  • 适用于需精准空间认知的AR/VR交互应用

现代扩展现实(XR)系统需对图像数据进行深度分析并融合传感器输入,要求增强现实(AR)/虚拟现实(VR)应用具备语义化3D场景推理能力。本文提出一种空间推理框架,将几何事实与符号谓词、关系相连接,解决3D物体间相对位置关系(如'在...上''在...后''靠近'等)的判定问题。其基础为定向3D边界框表示,并引入涵盖拓扑、连通性、方向性与朝向性的完整空间谓词集,采用接近自然语言的形式表达。生成的谓词构成空间知识图谱,结合基于流水线的推理模型,实现空间查询与动态规则评估。客户端与服务端实现表明,该框架能高效将几何数据转化为可操作知识,确保复杂3D环境中可扩展且技术无关的空间推理能力。该框架正推动空间本体构建,并与机器学习、自然语言处理及规则系统无缝集成,显著增强XR应用的智能水平。

原文摘要 · Abstract (English)

Modern extended reality XR systems provide rich analysis of image data and fusion of sensor input and demand AR/VR applications that can reason about 3D scenes in a semantic manner. We present a spatial reasoning framework that bridges geometric facts with symbolic predicates and relations to handle key tasks such as determining how 3D objects are arranged among each other ('on', 'behind', 'near', etc.). Its foundation relies on oriented 3D bounding box representations, enhanced by a comprehensive set of spatial predicates, ranging from topology and connectivity to directionality and orientation, expressed in a formalism related to natural language. The derived predicates form a spatial knowledge graph and, in combination with a pipeline-based inference model, enable spatial queries and dynamic rule evaluation. Implementations for client- and server-side processing demonstrate the framework's capability to efficiently translate geometric data into actionable knowledge, ensuring scalable and technology-independent spatial reasoning in complex 3D environments. The Spatial Reasoner framework is fostering the creation of spatial ontologies, and seamlessly integrates with and therefore enriches machine learning, natural language processing, and rule systems in XR applications.

空间推理XR应用3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。