arXiv:2509.01030cs.IR2025-09

用检索增强生成技术自动挖掘街道名称背后的历史来源。

Identifying Origins of Place Names via Retrieval Augmented Generation

  • 先从DBpedia中提取相关知识子图,再用微调语言模型排序
  • 在100个测试案例中准确率达78%,优于纯生成模型
  • 适合城市研究、历史地理与数字人文学者使用

街道名“蝙蝠侠街”背后的命名者是谁?理解地名背后的历史、文化与社会叙事,能揭示社区的深层脉络。尽管地名是地图集中的重要空间标识,但通常缺乏起源信息。当前补充地名来源依赖人工查阅海量文档,耗时费力。自然语言处理与语言模型的发展为自动化识别提供了可能。本文提出一种基于检索增强生成的流程,通过在大规模知识库DBpedia中搜索地名起源。给定空间查询后,系统首先提取可能相关的知识子图,再利用微调的语言模型(ColBERTv2和Llama2)对子图进行排序并生成答案。实验表明,语言模型常忽略文本中包含的空间信息,导致判断偏差。该方法揭示了地理信息检索中的关键挑战,并拓展了检索增强生成在地理领域的应用前景。

原文摘要 · Abstract (English)

Who is the "Batman" behind "Batman Street" in Melbourne? Understanding the historical, cultural, and societal narratives behind place names can reveal the rich context that has shaped a community. Although place names serve as essential spatial references in gazetteers, they often lack information about place name origins. Enriching these place names in today's gazetteers is a time-consuming, manual process that requires extensive exploration of a vast archive of documents and text sources. Recent advances in natural language processing and language models (LMs) hold the promise of significant automation of identifying place name origins due to their powerful capability to exploit the semantics of the stored documents. This chapter presents a retrieval augmented generation pipeline designed to search for place name origins over a broad knowledge base, DBpedia. Given a spatial query, our approach first extracts sub-graphs that may contain knowledge relevant to the query; then ranks the extracted sub-graphs to generate the final answer to the query using fine-tuned LM-based models (i.e., ColBERTv2 and Llama2). Our results highlight the key challenges facing automated retrieval of place name origins, especially the tendency of language models to under-use the spatial information contained in texts as a discriminating factor. Our approach also frames the wider implications for geographic information retrieval using retrieval augmented generation.

地名溯源检索增强知识图谱地理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。