用大模型让普通人也能安全查询城市地理数据,答得准还不会瞎编。
OGD4All: A Framework for Accessible Interaction with Geospatial Open Government Data Based on Large Language Models
- 通过语义检索+智能代码生成+沙盒执行,实现可追溯的交互式数据分析。
- 在199个问题上准确率达98%,召回率94%,能有效识别无数据支持的问题。
- 适合公众、政策制定者和研究人员使用,推动透明可信的开放治理。
我们提出OGD4All,一个基于大语言模型(LLMs)的透明、可审计、可复现框架,旨在提升公民对地理空间开放政府数据(OGD)的交互能力。系统融合语义数据检索、用于迭代代码生成的代理推理,以及安全沙盒执行,生成可验证的多模态输出。在涵盖430个苏黎世市数据集、11种LLM、199个问题的基准测试中,OGD4All达到98%的分析正确率和94%的召回率,同时可靠拒绝数据不支持的问题,显著降低幻觉风险。统计鲁棒性测试与专家反馈均证实其可靠性与社会价值。该方法展示了大模型如何提供可解释、多模态的公共数据访问,推动可信人工智能在开放治理中的应用。
原文摘要 · Abstract (English)
We present OGD4All, a transparent, auditable, and reproducible framework based on Large Language Models (LLMs) to enhance citizens' interaction with geospatial Open Government Data (OGD). The system combines semantic data retrieval, agentic reasoning for iterative code generation, and secure sandboxed execution that produces verifiable multimodal outputs. Evaluated on a 199-question benchmark covering both factual and unanswerable questions, across 430 City-of-Zurich datasets and 11 LLMs, OGD4All reaches 98% analytical correctness and 94% recall while reliably rejecting questions unsupported by available data, which minimizes hallucination risks. Statistical robustness tests, as well as expert feedback, show reliability and social relevance. The proposed approach shows how LLMs can provide explainable, multimodal access to public data, advancing trustworthy AI for open governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。