让普通人用自然语言查询城市3D模型,支持多轮对话和跨数据库查询。
CityLLM: A framework for natural-language querying of semantic 3D city models

- 基于大模型的流程,整合空间与图数据库实现自然语言交互。
- 54个查询测试中准确率最高达100%,平均重试少于3次。
- 适合城市规划、地理信息等非技术背景的研究者使用。
语义3D城市模型包含丰富的几何与语义信息,但因其复杂结构和专业格式,对非专家和跨学科研究者而言仍难访问与查询。为此,我们提出CityLLM框架,支持通过自然语言查询语义3D城市模型及互补的城市数据集。该框架结合空间与图数据库,在大模型驱动的工作流中支持多轮查询优化与跨数据库链式操作。我们在包含853栋LoD2建筑的鹿特丹CityJSON数据集上,使用GPT-OSS、Gemini 3.1和GPT-5.4及其变体进行评估,涵盖答案正确性、可视化正确性、查询成功率与重试次数等指标。共设计54个自然语言查询,覆盖空间、图谱、跨数据库与对话四种场景。结果表明,整体表现优异:答案正确率在85.2%至100%之间,可视化正确率达92.9%至100%,查询成功率为100%,所有查询平均重试次数少于3次。研究显示,CityLLM为语义3D城市数据提供了轻量且可扩展的对话式访问方式。
原文摘要 · Abstract (English)
Semantic 3D city models provide rich geometric and semantic information, but remain challenging for non-experts and interdisciplinary researchers to access and query due to their complex structures and specialized data formats. To address this issue, we present CityLLM, a framework for natural-language querying of semantic 3D city models alongside complementary urban datasets. The framework combines spatial and graph databases within an LLM-based workflow that supports iterative query refinement and cross-database chaining. We evaluate CityLLM on a CityJSON dataset of Rotterdam (853 LoD2 buildings) using GPT-OSS, Gemini 3.1, and GPT-5.4, along with selected variants, across multiple metrics: answer correctness, visualization correctness, query success, and retry attempts. A total of 54 natural-language queries are curated across four scenarios: spatial, graph, cross-database, and conversational. Results show strong overall performance, with answer correctness ranging from 85.2% to 100%, visualization correctness from 92.9% to 100%, a 100% query success rate, and fewer than three retries across all 54 queries. Overall, the findings suggest that CityLLM provides a lightweight and extensible approach for conversational access to semantic 3D city data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。