arXiv:2412.12961cs.CL2024-12

用大模型让普通人也能自然语言查询土地收购数据库。

Adaptations of AI models for querying the LandMatrix database in natural language

  • 用提示工程、RAG和智能体组合优化大模型查询能力。
  • 在LandMatrix数据集上验证,自然语言查询准确率显著提升。
  • 适合政策研究者、非技术人员快速获取土地数据。

Land Matrix(https://landmatrix.org)及其全球观察平台旨在为农业、资源开采和能源等领域的低收入与中等收入国家提供可靠的大型土地收购数据,以支持学术讨论和公共政策制定。尽管这些数据在学术界受到认可,但在公共政策领域仍被低估,主要原因是访问和使用复杂,需具备技术专长和对数据库结构的深入理解。本文旨在简化对不同数据库系统的数据访问。方法基于Land Matrix数据进行评估,对比多种大语言模型(LLMs)及提示工程、检索增强生成(RAG)、智能体等适配策略在查询GraphQL和REST接口时的表现。实验可复现,演示代码已开源:https://github.com/tetis-nlp/landmatrix-graphql-python。

原文摘要 · Abstract (English)

The Land Matrix initiative (https://landmatrix.org) and its global observatory aim to provide reliable data on large-scale land acquisitions to inform debates and actions in sectors such as agriculture, extraction, or energy in low- and middle-income countries. Although these data are recognized in the academic world, they remain underutilized in public policy, mainly due to the complexity of access and exploitation, which requires technical expertise and a good understanding of the database schema. The objective of this work is to simplify access to data from different database systems. The methods proposed in this article are evaluated using data from the Land Matrix. This work presents various comparisons of Large Language Models (LLMs) as well as combinations of LLM adaptations (Prompt Engineering, RAG, Agents) to query different database systems (GraphQL and REST queries). The experiments are reproducible, and a demonstration is available online: https://github.com/tetis-nlp/landmatrix-graphql-python.

自然语言查询大模型应用数据开放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。