arXiv:2411.17425cs.CV2024-11被引 1

用自监督视频分割提升历史地图地理实体对齐效果

Self-supervised Video Instance Segmentation Can Boost Geographic Entity Alignment in Historical Maps

  • 将历史地图实体分割与关联融合为统一视频分割任务
  • 自监督预训练使模型在无标注数据下提升24.9%的平均精度
  • 适合从事历史地理、文化遗产数字化的研究者

从历史地图中追踪地理实体(如建筑)可为文化传承、城市化变迁和环境演化研究提供关键信息。但跨图实体对齐仍具挑战,传统方法依赖两阶段流程:先检测再基于启发式匹配。本文提出一种结合视频实例分割(VIS)的新框架,实现端到端的实体分割与关联。由于高质量视频标注数据稀缺,尤其在包含数百至数千实体的历史地图上难以获取,我们探索自监督学习(SSL)以提升VIS性能。通过对比不同预训练配置,并设计一种从无标签历史地图图像生成合成视频的方法用于预训练,显著降低人工标注需求。实验表明,所提自监督VIS方法相比从零训练模型,平均精度(AP)提升24.9%,F1得分提高0.23。

原文摘要 · Abstract (English)

Tracking geographic entities from historical maps, such as buildings, offers valuable insights into cultural heritage, urbanization patterns, environmental changes, and various historical research endeavors. However, linking these entities across diverse maps remains a persistent challenge for researchers. Traditionally, this has been addressed through a two-step process: detecting entities within individual maps and then associating them via a heuristic-based post-processing step. In this paper, we propose a novel approach that combines segmentation and association of geographic entities in historical maps using video instance segmentation (VIS). This method significantly streamlines geographic entity alignment and enhances automation. However, acquiring high-quality, video-format training data for VIS models is prohibitively expensive, especially for historical maps that often contain hundreds or thousands of geographic entities. To mitigate this challenge, we explore self-supervised learning (SSL) techniques to enhance VIS performance on historical maps. We evaluate the performance of VIS models under different pretraining configurations and introduce a novel method for generating synthetic videos from unlabeled historical map images for pretraining. Our proposed self-supervised VIS method substantially reduces the need for manual annotation. Experimental results demonstrate the superiority of the proposed self-supervised VIS approach, achieving a 24.9\% improvement in AP and a 0.23 increase in F1 score compared to the model trained from scratch.

视频分割历史地图自监督学习地理实体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。