让增强现实理解物理空间语义,支持自然语言交互与实时学习。
Multimodal 3D Fusion and In-Situ Learning for Spatially Aware AI
- 融合视觉、语言与几何信息构建多模态3D对象表征
- 在Magic Leap 2上实现自然语言空间搜索与物品追踪
- 支持用户现场用语言指导系统学习物理环境变化
增强现实无缝融合虚拟与物理世界,依赖系统对物理环境的语义理解。本文提出一种统一语义、语言与几何信息的多模态3D对象表示,支持用户引导的机器学习。首先设计快速多模态3D重建流程,将CLIP视觉-语言特征融入环境与物体模型;随后提出“就地学习”机制,结合该表示,实现空间与语言双重意义的交互工具与界面。我们在Magic Leap 2上验证了两个真实应用:一是在物理环境中通过自然语言进行空间搜索;二是智能库存系统持续追踪物品随时间的变化。完整代码与演示数据已开源(https://github.com/cy-xu/spatially_aware_AI),推动空间感知智能研究。
原文摘要 · Abstract (English)
Seamless integration of virtual and physical worlds in augmented reality benefits from the system semantically "understanding" the physical environment. AR research has long focused on the potential of context awareness, demonstrating novel capabilities that leverage the semantics in the 3D environment for various object-level interactions. Meanwhile, the computer vision community has made leaps in neural vision-language understanding to enhance environment perception for autonomous tasks. In this work, we introduce a multimodal 3D object representation that unifies both semantic and linguistic knowledge with the geometric representation, enabling user-guided machine learning involving physical objects. We first present a fast multimodal 3D reconstruction pipeline that brings linguistic understanding to AR by fusing CLIP vision-language features into the environment and object models. We then propose "in-situ" machine learning, which, in conjunction with the multimodal representation, enables new tools and interfaces for users to interact with physical spaces and objects in a spatially and linguistically meaningful manner. We demonstrate the usefulness of the proposed system through two real-world AR applications on Magic Leap 2: a) spatial search in physical environments with natural language and b) an intelligent inventory system that tracks object changes over time. We also make our full implementation and demo data available at (https://github.com/cy-xu/spatially_aware_AI) to encourage further exploration and research in spatially aware AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。