用合成数据提升历史地图分割,自动模拟真实扫描噪声。
Automatic Uncertainty-Aware Synthetic Data Bootstrapping for Historical Map Segmentation
- 自动生成带视觉不确定性的合成历史地图,模仿真实扫描噪声。
- 在同质地图语料库上实现域自适应语义分割,提升模型性能。
- 适合历史地图分析、数字人文研究者,尤其缺标注数据时使用。
历史地图的自动化分析得益于深度学习在计算机视觉中的进展,但多数方法依赖大量标注数据,而这类数据对特定同质制图领域(即语料库)的历史地图通常难以获取。高质量训练数据的构建耗时且需大量人工。虽然合成数据可缓解真实样本稀缺问题,但常缺乏真实感和多样性。本文通过将历史地图语料库的制图风格迁移至现代矢量数据,自动生成无限量适用于土地覆盖解析等任务的合成历史地图。提出一种自动深度生成方法及替代性手动随机退化技术,以模拟历史地图扫描中常见的视觉不确定性(即主观不确定性)。为定量评估方法有效性,使用生成的数据集在同质地图语料库上进行域自适应语义分割,采用自构建图卷积网络进行评估,全面分析数据增强方法的影响。
原文摘要 · Abstract (English)
The automated analysis of historical documents, particularly maps, has drastically benefited from advances in deep learning and its success across various computer vision applications. However, most deep learning-based methods heavily rely on large amounts of annotated training data, which are typically unavailable for historical maps, especially for those belonging to specific, homogeneous cartographic domains, also known as corpora. Creating high-quality training data suitable for machine learning often takes a significant amount of time and involves extensive manual effort. While synthetic training data can alleviate the scarcity of real-world samples, it often lacks the affinity (realism) and diversity (variation) necessary for effective learning. By transferring the cartographic style of a historical map corpus onto modern vector data, we bootstrap an effectively unlimited number of synthetic historical maps suitable for tasks such as land-cover interpretation of a homogeneous historical map corpus. We propose an automatic deep generative approach and an alternative manual stochastic degradation technique to emulate the visual uncertainty and noise, also known as aleatoric uncertainty, commonly observed in historical map scans. To quantitatively evaluate the effectiveness and applicability of our approach, the bootstrapped training datasets were employed for domain-adaptive semantic segmentation on a homogeneous map corpus using a Self-Constructing Graph Convolutional Network, enabling a comprehensive assessment of the impact of our data bootstrapping methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。