用大模型理解地图和文本,自动定位生物标本采集地。
Large Multi-modal Model Cartographic Map Comprehension for Textual Locality Georeferencing
- 结合地图与文本的多模态方法,让模型理解空间关系。
- 零样本下平均定位误差约1公里,优于单一语言模型。
- 适合需要批量处理历史标本地理信息的科研人员。
过去几个世纪积累的数百万份生物样本记录存于自然史收藏中,但大多未进行地理标注。为这些复杂描述的采集地点进行地理标注是一项耗时费力的任务,现有自动化方法未能有效利用地图这一关键工具。本文提出一种新方法,利用大型多模态模型(LMM)的多模态能力,使模型能基于地图视觉上下文理解文本中的空间关系。采用网格化方法,将自回归模型适配至该任务的零样本设置。在小规模人工标注数据集上的实验显示,该方法平均距离误差约为1公里,显著优于仅使用大型语言模型的单模态地理标注及现有工具。论文还讨论了实验结果对LMM理解细粒度地图能力的启示,并提出了一个实用框架,可将该方法集成至地理标注工作流中。
原文摘要 · Abstract (English)
Millions of biological sample records collected in the last few centuries archived in natural history collections are un-georeferenced. Georeferencing complex locality descriptions associated with these collection samples is a highly labour-intensive task collection agencies struggle with. None of the existing automated methods exploit maps that are an essential tool for georeferencing complex relations. We present preliminary experiments and results of a novel method that exploits multi-modal capabilities of recent Large Multi-Modal Models (LMM). This method enables the model to visually contextualize spatial relations it reads in the locality description. We use a grid-based approach to adapt these auto-regressive models for this task in a zero-shot setting. Our experiments conducted on a small manually annotated dataset show impressive results for our approach ($\sim$1 km Average distance error) compared to uni-modal georeferencing with Large Language Models and existing georeferencing tools. The paper also discusses the findings of the experiments in light of an LMM's ability to comprehend fine-grained maps. Motivated by these results, a practical framework is proposed to integrate this method into a georeferencing workflow.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。