让大模型直接理解地理嵌入,提升空间推理效率与精度
Enabling Intrinsic Reasoning over Dense Geospatial Embeddings with DFR-Gemma
- 用轻量投影器将地理嵌入对齐大模型语义空间,直接注入为语义标记
- 零样本下在多任务基准上实现精准空间推理,效率显著优于文本基线
- 适合需要高效空间智能的科研与应用开发者,尤其关注地理数据分析
地理与时空数据的表征学习在通用地理智能中至关重要。近期的地理基础模型(如人口动态基础模型,PDFM)将复杂的人口与移动动态编码为紧凑嵌入。然而,其与大语言模型(LLM)的集成仍受限。现有方法将嵌入视为检索索引或转换为文本描述进行推理,导致冗余、分词效率低及数值误差。我们提出直接特征推理-Gemma(DFR-Gemma),一种新框架,使LLM可直接对密集地理嵌入进行推理。DFR通过轻量级投影器将高维嵌入与LLM潜在空间对齐,使嵌入作为语义标记与自然语言指令一同输入。该设计无需中间文本表示,实现对空间特征的内在推理。为评估此范式,我们引入一个包含多种问答任务的多任务地理基准,涵盖特征查询、比较与语义描述。实验表明,DFR使LLM能解码潜在空间模式,在跨任务上实现准确的零样本推理,同时相比文本基线大幅提升效率。结果表明,将嵌入作为主要输入数据,为多模态地理智能提供了更直接、高效且可扩展的路径。
原文摘要 · Abstract (English)
Representation learning for geospatial and spatio-temporal data plays a critical role in enabling general-purpose geospatial intelligence. Recent geospatial foundation models, such as the Population Dynamics Foundation Model (PDFM), encode complex population and mobility dynamics into compact embeddings. However, their integration with Large Language Models (LLMs) remains limited. Existing approaches to LLM integration treat these embeddings as retrieval indices or convert them into textual descriptions for reasoning, introducing redundancy, token inefficiency, and numerical inaccuracies. We propose Direct Feature Reasoning-Gemma (DFR-Gemma), a novel framework that enables LLMs to reason directly over dense geospatial embeddings. DFR aligns high-dimensional embeddings with the latent space of an LLM via a lightweight projector, allowing embeddings to be injected as semantic tokens alongside natural language instructions. This design eliminates the need for intermediate textual representations and enables intrinsic reasoning over spatial features. To evaluate this paradigm, we introduce a multi-task geospatial benchmark that pairs embeddings with diverse question-answer tasks, including feature querying, comparison, and semantic description. Experimental results show that DFR allows LLMs to decode latent spatial patterns and perform accurate zero-shot reasoning across tasks, while significantly improving efficiency compared to text-based baselines. Our results demonstrate that treating embeddings as primary data inputs, provides a more direct, efficient, and scalable approach to multimodal geospatial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。