自动识别历史地图中的多词地名,提升古地图检索效率
Automatic Search of Multiword Place Names on Historical Maps
- 通过最小生成树连接空间相近、格式相似的单字标签,构建多词地名候选
- 实验验证该方法能准确识别多词地名,有效召回大量跨时代的相关地图
- 适合研究历史地理变迁的学者,尤其适用于多词地名密集的历史地图库
历史地图是了解过去的重要信息来源,数字化后的扫描地图在在线图书馆中日益普及。为从海量地图中检索包含特定地点的地图,已有研究利用计算机视觉技术识别地图上的文字,实现对单个地名的搜索。但多词地名因文本布局复杂而难以识别。本文提出一种高效查询方法,用于在历史地图上搜索给定的多词地名。基于现有文字识别结果,通过构建最小生成树将单字文本标签关联为潜在的多词短语,目标是连接空间接近、高度、角度和大小写相似的标签对。随后以这些树结构为单位进行查询。我们设计两个实验:1)评估最小生成树方法在连接多词地名方面的准确性;2)评估查询方法可检索到的地图数量及其时间跨度。结果揭示了使用多词名称的地点在大量历史地图中的演变情况。
原文摘要 · Abstract (English)
Historical maps are invaluable sources of information about the past, and scanned historical maps are increasingly accessible in online libraries. To retrieve maps from these large libraries that contain specific places of interest, previous work has applied computer vision techniques to recognize words on historical maps, enabling searches for maps that contain specific place names. However, searching for multiword place names is challenging due to complex layouts of text labels on historical maps. This paper proposes an efficient query method for searching a given multiword place name on historical maps. Using existing methods to recognize words on historical maps, we link single-word text labels into potential multiword phrases by constructing minimum spanning trees. These trees aim to link pairs of text labels that are spatially close and have similar height, angle, and capitalization. We then query these trees for the given multiword place name. We evaluate the proposed method in two experiments: 1) to evaluate the accuracy of the minimum spanning tree approach at linking multiword place names and 2) to evaluate the number and time range of maps retrieved by the query approach. The resulting maps reveal how places using multiword names have changed on a large number of maps from across history.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。