首个评估视觉语言模型解读规划地图能力的基准,揭示AI在专业决策中的短板。
PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

- 构建223张专家标注的规划地图数据集,涵盖1629个问答对。
- 模型在实施类任务上仍差强人意,最优模型比前代提升27%但未达人类水平。
- 适合关注城市规划、多模态智能评估的研究者与政策科技从业者。
空间规划图是国土治理的核心,将规划目标、法规与空间策略转化为可视形式,用于决策、公众沟通与机构协调。其解读需精细视觉感知、空间推理与政策导向的专业判断,对人类学习者和AI系统均构成挑战。尽管视觉-语言模型(VLMs)在城市规划分析中日益受到关注,现有多模态基准多聚焦通用视觉理解,忽视了规划实践中的领域特定认知过程。为此,我们提出PlanBench-V,首个面向视觉语言模型在空间规划地图解读能力的综合性评估基准。首先构建了由专业规划师标注的《空间规划地图数据库》(SPMD),包含223张规划图与1629个问答对,覆盖多样地理区域与制图风格。随后提出基于理论的评估框架,从感知、推理、关联到实施四个递进维度衡量模型能力。在两代VLM上的大量实验显示,虽有进展但仍存在明显局限:最佳2026年代理推理模型Qwen3.6-Plus相较2025年最佳模型GPT-4o提升27%,但在需要评价性判断、政策敏感性和约束意识的实施类任务上表现仍弱。这些发现揭示当前VLM在专业规划场景下的根本缺陷,凸显开发领域自适应多模态推理框架的必要性。代码与数据已开源于https://plangpt.github.io。
原文摘要 · Abstract (English)
Spatial planning maps are central to territorial governance, translating planning objectives, regulations, and spatial strategies into visual forms for decision-making, public communication, and institutional coordination. Their interpretation, however, requires fine-grained visual perception, spatial reasoning, and policy-informed professional judgment, creating major challenges for both human learners and AI systems. With the rapid progress of Vision-Language Models (VLMs), their use in urban planning analysis is gaining attention, yet existing multimodal benchmarks mainly target general visual understanding and overlook the domain-specific cognitive processes of planning practice. To address this gap, we introduce PlanBench-V, the first comprehensive benchmark for evaluating VLMs in spatial planning map interpretation. We first build the Spatial Planning Map Database (SPMD), an expert-annotated dataset of 223 planning maps and 1629 question-answer pairs curated by professional planners, covering diverse geographic regions and cartographic styles. We then propose a theory-informed evaluation framework assessing four progressive capabilities: Perception, Reasoning, Association, and Implementation, corresponding to the cognitive pipeline of planning map interpretation. Extensive experiments across two generations of VLMs show clear progress but persistent limitations. The best 2026 agentic reasoning model, Qwen3.6-Plus, substantially outperforms the best 2025 model, GPT-4o, by 27%. Nevertheless, all models still struggle with implementation-oriented tasks requiring evaluative judgment, policy sensitivity, and constraint-aware decision-making. These findings reveal fundamental limitations of current VLMs in professional planning contexts and highlight the need for domain-adaptive multimodal reasoning frameworks. Code and data are available at https://plangpt.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。