arXiv:2510.02097cs.CV2025-10

用AI从历史地图提取法国1925-1950年城市范围,首份全国开放数据集。

Mapping Historic Urban Footprints in France: Balancing Quality, Scalability and AI Techniques

  • 双阶段U-Net处理历史地图的复杂纹理与辐射差异。
  • 整体准确率达73%,覆盖全法941张高分辨率图幅。
  • 适合研究长期城市化、历史地理与文化遗产的学者。

1970年前法国历史城市扩张的定量分析受限于缺乏全国性的数字城市边界数据。本研究通过构建可扩展的深度学习流水线,从1925-1950年的Scan Histo历史地图系列中提取城市区域,首次生成该关键时期的开放获取、全国尺度城市边界数据集。核心创新为双阶段U-Net方法,第一阶段在初始数据集上训练,生成初步地图以识别文本、道路等混淆区域,指导针对性数据增强;第二阶段使用优化数据集及第一阶段二值输出,有效抑制辐射噪声,显著降低误检。在高性能计算集群上部署,处理覆盖整个本土法国的941张高分辨率图幅。最终拼接成果整体准确率达73%,有效捕捉多样化城市形态,克服标签与等高线等常见伪影。代码、训练数据与生成的全国城市栅格数据均公开发布,支持长期城市化动态研究。

原文摘要 · Abstract (English)

Quantitative analysis of historical urban sprawl in France before the 1970s is hindered by the lack of nationwide digital urban footprint data. This study bridges this gap by developing a scalable deep learning pipeline to extract urban areas from the Scan Histo historical map series (1925-1950), which produces the first open-access, national-scale urban footprint dataset for this pivotal period. Our key innovation is a dual-pass U-Net approach designed to handle the high radiometric and stylistic complexity of historical maps. The first pass, trained on an initial dataset, generates a preliminary map that identifies areas of confusion, such as text and roads, to guide targeted data augmentation. The second pass uses a refined dataset and the binarized output of the first model to minimize radiometric noise, which significantly reduces false positives. Deployed on a high-performance computing cluster, our method processes 941 high-resolution tiles covering the entirety of metropolitan France. The final mosaic achieves an overall accuracy of 73%, effectively capturing diverse urban patterns while overcoming common artifacts like labels and contour lines. We openly release the code, training datasets, and the resulting nationwide urban raster to support future research in long-term urbanization dynamics.

城市化历史地图深度学习地理信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。