测试大模型在图像与文本间传递空间信息的能力,发现其仍存在明显瓶颈。
LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents

- 设计图像到文本、再从文本还原图像的双阶段任务,评估模态转移能力。
- 当前主流大模型(如OpenAI)在还原简单色块网格时准确率不足,表现不佳。
- 揭示多模态对齐是实现地理空间智能的关键挑战,适合研究多模态模型的学者参考。
AI模型在理解与处理空间信息方面日益成熟,推动了空间任务中的智能体式问题求解。然而,现有研究大多仅以文本作为输入输出模态,与人类在地理信息系统(GIS)工作流中同时、交替使用文本与视觉信息的方式不符。为真正实现自动化地理空间分析流程或复现人类设计的工作流,大型多模态模型(LMMs)需具备在图像与文本模态间无缝切换的能力。本文提出一种模态转移任务:首先让一个LMM描述一张由彩色方块组成的规则网格图像;随后,另一个LMM实例根据该文本描述重建原始空间场景的图像。该任务量化了LMM在图像与文本间传递空间信息的能力。通过空间信息理论视角分析发现,当前先进LMM(如OpenAI系列)在还原简单色块网格图像时仍表现不佳,暴露出多模态对齐不足这一关键瓶颈,表明实现鲁棒的地理空间理解亟需更严格的跨模态对齐。
原文摘要 · Abstract (English)
AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial tasks and workflows. However, most of the research on their spatial capabilities (e.g., spatial reasoning) has focused on the textual modality as input and output. This contrasts with the human approach to GIS workflows, where text and visual modalities are often used together, interchangeably, and in a complementary manner. Thus, to truly achieve an automated GIS analysis pipeline or carry out human-designed GIS workflows, AI models --- Large Multimodal Models (LMMs) in particular --- need to be able to seamlessly transition between image- and text-based modalities that are traditionally used in such workflows. We present a modality transfer task that (1) asks an LMM to first describe an input image of colored squares in a regular grid, and (2) asks a new LMM instance to re-generate an image of the original spatial scene using the textual description output by the former model. This task quantifies the ability of LMMs to transfer spatial information between image and text modalities. Ultimately, by examining the modality transfer capability of LMMs through the lens of spatial information theory, this work highlights a critical bottleneck: achieving strong and robust geospatial understanding in LMMs requires rigorous, multi-modal alignment. Our results indicate that recent LMMs (here from OpenAI) still struggle with modality transfer, when tasked with re-generating an image of a simple spatial grid of color squares.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。