arXiv:2510.08385cs.CVcs.AI2025-10被引 1

用GPT-4o+上下文学习,自动识别古地图图例并配对符号与文字。

Detecting Legend Items on Historical Maps Using GPT-4o with In-Context Learning

  • 结合LayoutLMv3与GPT-4o,通过边界框预测实现图例项与描述的匹配。
  • 在测试集上达到88%的F-1和85%的IoU,优于基线方法。
  • 适合需要高效解析多风格古地图的数字人文与档案研究者。

历史地图的图例对于解读地图符号至关重要。然而,其布局不一致且格式非结构化,导致自动提取困难。以往工作主要集中在分割或通用光学字符识别(OCR),缺乏有效将图例符号与其对应描述结构化匹配的方法。本文提出一种结合LayoutLMv3进行版面检测,并利用GPT-4o的上下文学习能力,通过结构化JSON提示完成图例项与描述的检测与链接。实验表明,使用结构化JSON提示的GPT-4在测试中达到88%的F-1和85%的交并比(IoU),优于基线方法;同时揭示了提示设计、示例数量及版面对齐对性能的影响。该方法支持可扩展的、版面感知的图例解析,显著提升不同视觉风格历史地图的索引与可检索性。

原文摘要 · Abstract (English)

Historical map legends are critical for interpreting cartographic symbols. However, their inconsistent layouts and unstructured formats make automatic extraction challenging. Prior work focuses primarily on segmentation or general optical character recognition (OCR), with few methods effectively matching legend symbols to their corresponding descriptions in a structured manner. We present a method that combines LayoutLMv3 for layout detection with GPT-4o using in-context learning to detect and link legend items and their descriptions via bounding box predictions. Our experiments show that GPT-4 with structured JSON prompts outperforms the baseline, achieving 88% F-1 and 85% IoU, and reveal how prompt design, example counts, and layout alignment affect performance. This approach supports scalable, layout-aware legend parsing and improves the indexing and searchability of historical maps across various visual styles.

图例识别GPT-4o古地图结构化提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。