arXiv:2512.00294cs.CVcs.AI2025-12

让AR系统听懂复杂自然语言指令,精准定位物体并理解空间关系。

Words into World: A Task-Adaptive Agent for Language-Guided Spatial Retrieval in AR

  • 用多模态大模型+视觉定位工具,动态响应不同难度的语言查询。
  • 能返回米级精度3D坐标,支持多物体关系推理与交互操作。
  • 适合需要精准空间理解的AR应用,如智能导航、人机协作。

传统增强现实系统依赖固定类别检测器或标记物,难以理解复杂的开放词汇自然语言查询。本文提出一种模块化AR代理系统,融合多模态大语言模型(MLLMs)与具身视觉模型,实现空间与语言的关联推理及语言引导的空间检索。该自适应任务代理协调MLLMs与坐标感知感知工具,应对从简单物体识别到多物体关系推理等不同复杂度的查询,并返回米级精度的3D锚点。系统构建动态AR场景图,编码九种类型关系(空间、结构语义、因果功能),使MLLMs不仅能识别物体,还能理解其在三维空间中的关联与互动。通过任务自适应的感兴趣区域高亮和上下文空间检索,系统引导人类注意力至信息密集区,支持人机协同优化。针对复杂查询,代理动态调用坐标感知工具(如选择、测量、比较、操作),将语言理解落地于物理操作。模块化设计支持即插即用视觉语言模型,无需重训练,使AR代理成为连接大模型与真实世界空间智能的桥梁。此外,我们提出GroundedAR-Bench,一个评估语言驱动现实定位与关系具身化的基准框架,覆盖多样化环境。

原文摘要 · Abstract (English)

Traditional augmented reality (AR) systems predominantly rely on fixed class detectors or fiducial markers, limiting their ability to interpret complex, open-vocabulary natural language queries. We present a modular AR agent system that integrates multimodal large language models (MLLMs) with grounded vision models to enable relational reasoning in space and language-conditioned spatial retrieval in physical environments. Our adaptive task agent coordinates MLLMs and coordinate-aware perception tools to address varying query complexities, ranging from simple object identification to multi-object relational reasoning, while returning meter-accurate 3D anchors. It constructs dynamic AR scene graphs encoding nine typed relations (spatial, structural-semantic, causal-functional), enabling MLLMs to understand not just what objects exist, but how they relate and interact in 3D space. Through task-adaptive region-of-interest highlighting and contextual spatial retrieval, the system guides human attention to information-dense areas while supporting human-in-the-loop refinement. The agent dynamically invokes coordinate-aware tools for complex queries-selection, measurement, comparison, and actuation-grounding language understanding in physical operations. The modular architecture supports plug-and-use vision-language models without retraining, establishing AR agents as intermediaries that augment MLLMs with real-world spatial intelligence for interactive scene understanding. We also introduce GroundedAR-Bench, an evaluation framework for language-driven real world localization and relation grounding across diverse environments.

AR空间推理大模型语言导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。