用合成数据提升大模型在地图上追踪路径的能力。
MapTrace: Scalable Data Generation for Route Tracing on Maps
- 通过合成地图和像素级解析自动生成精确标注数据。
- 构建2.3万条路径样本,使模型成功率最高提升6.4点。
- 适合研究地图理解、空间推理与合成数据生成的学者。
尽管多模态大模型在诸多视觉与文本推理任务中已达到人类水平,但在细粒度空间理解(如地图路径追踪)方面仍表现不足。与人类能快速解析地图不同,现有模型常忽视基本路径约束,部分原因在于收集大规模、像素级路径标注成本过高且难度大。为此,我们提出一种可扩展的合成数据生成流程,利用合成地图图像与像素级解析,自动生成该任务的精准标注。基于此流程,我们构建了包含2.3万条路径样本、覆盖4000张地图的微调数据集,使模型获得更接近人类的空间能力。在此数据集上微调开源与专有多模态大模型后,MapBench测试显示,模型鲁棒性显著提升,成功率最高提高6.4个百分点,同时路径追踪误差(NDTW)下降。结果表明,预训练模型中缺失的细粒度空间推理能力可通过合成监督显式教学。
原文摘要 · Abstract (English)
While Multimodal Large Language Models have achieved human-like performance on many visual and textual reasoning tasks, their proficiency in fine-grained spatial understanding, such as route tracing on maps remains limited. Unlike humans, who can quickly learn to parse and navigate maps, current models often fail to respect fundamental path constraints, in part due to the prohibitive cost and difficulty of collecting large-scale, pixel-accurate path annotations. To address this, we introduce a scalable synthetic data generation pipeline that leverages synthetic map images and pixel-level parsing to automatically produce precise annotations for this challenging task. Using this pipeline, we construct a fine-tuning dataset of 23k path samples across 4k maps, enabling models to acquire more human-like spatial capabilities. Using this dataset, we fine-tune both open-source and proprietary MLLMs. Results on MapBench show that finetuning substantially improves robustness, raising success rates by up to 6.4 points, while also reducing path-tracing error (NDTW). These gains highlight that fine-grained spatial reasoning, absent in pretrained models, can be explicitly taught with synthetic supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。