arXiv:2506.23352cs.CV2025-06ICCV被引 2

用自然语言操作城市级3D场景,支持复杂地理推理。

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

  • 构建地理感知的3D语言场与视觉工具链,实现大规模场景理解。
  • 在952个任务上超越现有模型,尤其在空间推理和计数上提升显著。
  • 适合城市规划、自动驾驶等需要复杂地理交互的领域使用。

3D语言场的发展使用户可通过自然语言与3D场景互动,但现有方法多局限于小规模环境,缺乏对大型复杂城市场景的可扩展性与组合推理能力。为此,我们提出GeoProg3D,一个面向高保真城市级3D场景的自然语言交互视觉编程框架。该框架包含两个核心组件:(i) 地理感知的城市级3D语言场(GCLF),采用高效的分层3D建模结构,结合地理信息(方向、距离、高程、地标)实现对大规模城市空间的快速筛选;(ii) 地理视觉API(GV-APIs),包括区域分割与目标检测等专用视觉工具。框架利用大语言模型(LLMs)作为推理引擎,动态调用GV-APIs并操作GCLF,支持多样化的地理视觉任务。为评估城市级推理性能,我们构建了包含952个问答对的GeoEval3D基准数据集,涵盖定位、空间推理、比较、计数与测量五大挑战任务。实验表明,GeoProg3D在多个任务中显著优于现有3D语言场与视觉语言模型。据我们所知,这是首个通过自然语言实现高保真城市级3D环境中组合地理推理的框架。代码已公开于https://snskysk.github.io/GeoProg3D/。

原文摘要 · Abstract (English)

The advancement of 3D language fields has enabled intuitive interactions with 3D scenes via natural language. However, existing approaches are typically limited to small-scale environments, lacking the scalability and compositional reasoning capabilities necessary for large, complex urban settings. To overcome these limitations, we propose GeoProg3D, a visual programming framework that enables natural language-driven interactions with city-scale high-fidelity 3D scenes. GeoProg3D consists of two key components: (i) a Geography-aware City-scale 3D Language Field (GCLF) that leverages a memory-efficient hierarchical 3D model to handle large-scale data, integrated with geographic information for efficiently filtering vast urban spaces using directional cues, distance measurements, elevation data, and landmark references; and (ii) Geographical Vision APIs (GV-APIs), specialized geographic vision tools such as area segmentation and object detection. Our framework employs large language models (LLMs) as reasoning engines to dynamically combine GV-APIs and operate GCLF, effectively supporting diverse geographic vision tasks. To assess performance in city-scale reasoning, we introduce GeoEval3D, a comprehensive benchmark dataset containing 952 query-answer pairs across five challenging tasks: grounding, spatial reasoning, comparison, counting, and measurement. Experiments demonstrate that GeoProg3D significantly outperforms existing 3D language fields and vision-language models across multiple tasks. To our knowledge, GeoProg3D is the first framework enabling compositional geographic reasoning in high-fidelity city-scale 3D environments via natural language. The code is available at https://snskysk.github.io/GeoProg3D/.

3D语言场地理推理视觉编程城市建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。