arXiv:2508.08266cs.LGcs.AI2025-08

用大模型自动定位殖民地时期的弗吉尼亚土地契约坐标,准确率超人工。

Benchmarking Large Language Models for Geolocating Colonial Virginia Land Grants

  • 用大模型直接或调用工具链将土地契约文字转为地理坐标。
  • 最优模型平均误差23公里,比最差模型快53.5%,比人工快67%。
  • 适合历史学者、数字人文研究者做大规模历史地理分析。

17至18世纪弗吉尼亚的土地专利主要以叙事性的测距边界描述留存,限制了空间分析。本研究系统评估当前大语言模型(LLMs)在将这些文本描述转换为精确经纬度坐标方面的表现。释放了一个包含5,471份弗吉尼亚专利摘要(1695–1732年)的数字化语料库,并构建了43个经严格验证的测试案例作为初始地理聚焦基准。测试了六款OpenAI模型(o系列、GPT-4类、GPT-3.5类),采用直接输出坐标和调用外部地理编码API的链式思维两种范式。结果与GIS分析师基准、Stanford NER地理标注器、Mordecai-3神经地理标注器及县中心启发式方法对比。最优单次调用模型o3-2025-04-16平均误差23公里(中位数14公里),优于中位数LLM(37.4公里)37.5%,劣于最弱模型(50.3公里)53.5%,优于人工(67%)、Stanford NER(70%)。五次调用集成进一步将误差降至19.2公里(中位数12.2公里),成本仅增加约0.20美元/契约,优于中位数模型48.7%。删除申请人姓名后误差略有上升(约7%),表明模型依赖文本中的地标和邻近描述而非记忆。成本效益高的gpt-4o-2024-08-06模型保持28公里平均误差,每千份仅需1.09美元,确立了高性价比基准。外部地理编码工具在此任务中未带来显著提升。结果表明,大模型在可扩展、高精度、低成本的历史地理标注方面具有巨大潜力。

原文摘要 · Abstract (English)

Virginia's seventeenth- and eighteenth-century land patents survive primarily as narrative metes-and-bounds descriptions, limiting spatial analysis. This study systematically evaluates current-generation large language models (LLMs) in converting these prose abstracts into geographically accurate latitude/longitude coordinates within a focused evaluation context. A digitized corpus of 5,471 Virginia patent abstracts (1695-1732) is released, with 43 rigorously verified test cases serving as an initial, geographically focused benchmark. Six OpenAI models across three architectures-o-series, GPT-4-class, and GPT-3.5-were tested under two paradigms: direct-to-coordinate and tool-augmented chain-of-thought invoking external geocoding APIs. Results were compared against a GIS analyst baseline, Stanford NER geoparser, Mordecai-3 neural geoparser, and a county-centroid heuristic. The top single-call model, o3-2025-04-16, achieved a mean error of 23 km (median 14 km), outperforming the median LLM (37.4 km) by 37.5%, the weakest LLM (50.3 km) by 53.5%, and external baselines by 67% (GIS analyst) and 70% (Stanford NER). A five-call ensemble further reduced errors to 19.2 km (median 12.2 km) at minimal additional cost (~USD 0.20 per grant), outperforming the median LLM by 48.7%. A patentee-name redaction ablation slightly increased error (~7%), showing reliance on textual landmark and adjacency descriptions rather than memorization. The cost-effective gpt-4o-2024-08-06 model maintained a 28 km mean error at USD 1.09 per 1,000 grants, establishing a strong cost-accuracy benchmark. External geocoding tools offer no measurable benefit in this evaluation. These findings demonstrate LLMs' potential for scalable, accurate, cost-effective historical georeferencing.

历史地理大模型应用地理编码数字人文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。