arXiv:2508.06584cs.DBcs.AI2025-08

提出统一编码多种地理形状的模型,提升空间实体匹配精度。

Omni Geometry Representation Learning vs Large Language Models for Geospatial Entity Resolution

  • 设计全能几何编码器,支持点、线、多边形等多类地理形状嵌入
  • 在新数据集上比现有方法提升12%的F1值,显著优于简化几何的旧方法
  • 验证大模型提示策略的有效性,为未来空间匹配提供新思路

地理空间数据库的构建与维护依赖高效的实体匹配技术。尽管兴趣点(POI)匹配已广泛研究,但具有多样化几何形态的实体匹配仍被忽视,主要因缺乏将异构几何体无缝嵌入神经网络的统一方法。现有方法常将复杂几何简化为单一点,造成大量空间信息丢失。为此,我们提出Omni模型,其配备全能几何编码器,可嵌入点、线、折线、多边形及复合多边形等多种几何类型,有效捕捉空间实体的复杂特征。同时,Omni采用基于Transformer的预训练语言模型,在个体文本属性上通过属性亲和机制进行建模。模型在现有仅含点的数据集及新构建的多样化几何地理空间实体匹配数据集上进行了严格测试。结果表明,Omni相比现有方法提升最高达12%(F1),表现显著。此外,我们探索了大语言模型(LLM)在地理空间实体匹配中的潜力,通过不同提示策略与学习场景对比,结果显示大语言模型具备竞争力。

原文摘要 · Abstract (English)

The development, integration, and maintenance of geospatial databases rely heavily on efficient and accurate matching procedures of Geospatial Entity Resolution (ER). While resolution of points-of-interest (POIs) has been widely addressed, resolution of entities with diverse geometries has been largely overlooked. This is partly due to the lack of a uniform technique for embedding heterogeneous geometries seamlessly into a neural network framework. Existing neural approaches simplify complex geometries to a single point, resulting in significant loss of spatial information. To address this limitation, we propose Omni, a geospatial ER model featuring an omni-geometry encoder. This encoder is capable of embedding point, line, polyline, polygon, and multi-polygon geometries, enabling the model to capture the complex geospatial intricacies of the places being compared. Furthermore, Omni leverages transformer-based pre-trained language models over individual textual attributes of place records in an Attribute Affinity mechanism. The model is rigorously tested on existing point-only datasets and a new diverse-geometry geospatial ER dataset. Omni produces up to 12% (F1) improvement over existing methods. Furthermore, we test the potential of Large Language Models (LLMs) to conduct geospatial ER, experimenting with prompting strategies and learning scenarios, comparing the results of pre-trained language model-based methods with LLMs. Results indicate that LLMs show competitive results.

地理实体匹配几何编码大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。