专为城市规划图设计的视觉语言模型,提升地图分析准确性与效率。
PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models
- 基于领域数据构建高质量问答训练集,精准匹配规划图理解需求。
- 通过结构化验证机制降低幻觉,保持高事实准确性。
- 轻量7B模型性能媲美超72B大模型,适合实际部署与教学使用。
在城市规划领域,现有视觉语言模型(VLMs)难以有效分析和评估规划地图,而这些视觉要素对规划师及教育场景至关重要。规划地图涉及土地利用、基础设施布局与功能分区,需具备空间配置、法规要求及多尺度分析的专业理解。为此,我们提出首个面向城市规划地图的专用视觉语言模型PlanGPT-VL,采用三项创新:(1) 利用PlanAnno-V框架合成高质量视觉问答数据;(2) 引入关键点思维机制,通过结构化验证减少幻觉;(3) 结合监督微调与冻结视觉编码器的综合训练方法。在自建的PlanBench-V基准上系统评估显示,PlanGPT-VL在专业地图解析任务中显著优于通用顶尖VLMs,为规划专业人士提供可靠的地图分析、评估与教育工具,同时保持高事实准确性。其轻量级7B参数模型表现可比超过72B参数的模型,证明了高效领域专业化无需牺牲性能。
原文摘要 · Abstract (English)
In the field of urban planning, existing Vision-Language Models (VLMs) frequently fail to effectively analyze and evaluate planning maps, despite the critical importance of these visual elements for urban planners and related educational contexts. Planning maps, which visualize land use, infrastructure layouts, and functional zoning, require specialized understanding of spatial configurations, regulatory requirements, and multi-scale analysis. To address this challenge, we introduce PlanGPT-VL, the first domain-specific Vision-Language Model tailored specifically for urban planning maps. PlanGPT-VL employs three innovative approaches: (1) PlanAnno-V framework for high-quality VQA data synthesis, (2) Critical Point Thinking to reduce hallucinations through structured verification, and (3) comprehensive training methodology combining Supervised Fine-Tuning with frozen vision encoder parameters. Through systematic evaluation on our proposed PlanBench-V benchmark, we demonstrate that PlanGPT-VL significantly outperforms general-purpose state-of-the-art VLMs in specialized planning map interpretation tasks, offering urban planning professionals a reliable tool for map analysis, assessment, and educational applications while maintaining high factual accuracy. Our lightweight 7B parameter model achieves comparable performance to models exceeding 72B parameters, demonstrating efficient domain specialization without sacrificing performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。