用视频思维自动分割历史地图,提升时空信息提取效率
MapSAM2: Adapting SAM2 for Automatic Segmentation of Historical Map Images and Time Series
- 将地图图像和时序数据都视为视频,利用上下文记忆增强几何精度
- 在少样本条件下实现建筑、区域等特征的准确分割与时间关联
- 提出伪时序生成法降低标注成本,适合地理分析与文化遗产研究
历史地图是记录不同时期地理特征的独特宝贵资料,但其风格差异大、标注数据稀缺,自动化分析仍具挑战。从历史地图时序中构建关联的时空数据集尤其耗时费力,需整合多幅地图信息。此类数据对建筑年代判定、道路网络演化、聚落发展及环境变迁研究至关重要。我们提出MapSAM2,一个统一框架,用于自动分割历史地图图像与时序数据。基于视觉基础模型,通过少样本微调适应多种分割任务。核心创新在于将地图图像与时序数据均视为视频:对单幅图像,将多个图块处理为视频,利用记忆注意力机制融合相似图块的上下文线索,显著提升面状要素的几何准确性;对时序数据,我们引入标注的Siegfried Building Time Series Dataset,为降低标注成本,提出从单年份地图模拟常见时变变换生成伪时序数据。实验表明,MapSAM2能有效学习时间关联,在有限监督或伪视频条件下准确分割并链接建筑物。我们将公开数据集与代码以支持后续研究。
原文摘要 · Abstract (English)
Historical maps are unique and valuable archives that document geographic features across different time periods. However, automated analysis of historical map images remains a significant challenge due to their wide stylistic variability and the scarcity of annotated training data. Constructing linked spatio-temporal datasets from historical map time series is even more time-consuming and labor-intensive, as it requires synthesizing information from multiple maps. Such datasets are essential for applications such as dating buildings, analyzing the development of road networks and settlements, studying environmental changes etc. We present MapSAM2, a unified framework for automatically segmenting both historical map images and time series. Built on a visual foundation model, MapSAM2 adapts to diverse segmentation tasks with few-shot fine-tuning. Our key innovation is to treat both historical map images and time series as videos. For images, we process a set of tiles as a video, enabling the memory attention mechanism to incorporate contextual cues from similar tiles, leading to improved geometric accuracy, particularly for areal features. For time series, we introduce the annotated Siegfried Building Time Series Dataset and, to reduce annotation costs, propose generating pseudo time series from single-year maps by simulating common temporal transformations. Experimental results show that MapSAM2 learns temporal associations effectively and can accurately segment and link buildings in time series under limited supervision or using pseudo videos. We will release both our dataset and code to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。