arXiv:2504.21027cs.CLcs.AI2025-04被引 5

评测大模型在城市规划中的表现,发现其法规理解能力严重不足。

UrbanPlanBench: A Comprehensive Urban Planning Benchmark for Evaluating Large Language Models

  • 构建城市规划专用评测基准UrbanPlanBench,覆盖专业知识与法规。
  • 70%的大模型在规划法规理解上表现不佳,远低于专业标准。
  • 提供超3万条指令数据集,助力模型提升规划领域能力。

大型语言模型(LLMs)有望变革依赖人类专家经验的传统领域。城市规划作为塑造日常环境的关键专业,高度依赖多维度知识与实践经验。当前大模型在该领域的辅助能力仍不明确。本文提出综合性评测基准UrbanPlanBench,涵盖基础原理、专业知识、管理与法规,贴近人类规划师资质要求。评估显示,大模型在规划知识获取上存在显著不平衡,即使最先进的模型也未达专业标准。例如,70%的模型在理解规划法规方面表现欠佳。此外,我们发布了迄今最大的监督微调数据集UrbanPlanText,包含超3万条来自规划考试与教材的指令对。微调后模型在记忆测试与知识理解上有所提升,但在需领域术语与推理的任务中仍有巨大改进空间。相关基准、数据集及工具链已开源,旨在推动大模型与人类专家在城市规划中的协同创新。

原文摘要 · Abstract (English)

The advent of Large Language Models (LLMs) holds promise for revolutionizing various fields traditionally dominated by human expertise. Urban planning, a professional discipline that fundamentally shapes our daily surroundings, is one such field heavily relying on multifaceted domain knowledge and experience of human experts. The extent to which LLMs can assist human practitioners in urban planning remains largely unexplored. In this paper, we introduce a comprehensive benchmark, UrbanPlanBench, tailored to evaluate the efficacy of LLMs in urban planning, which encompasses fundamental principles, professional knowledge, and management and regulations, aligning closely with the qualifications expected of human planners. Through extensive evaluation, we reveal a significant imbalance in the acquisition of planning knowledge among LLMs, with even the most proficient models falling short of meeting professional standards. For instance, we observe that 70% of LLMs achieve subpar performance in understanding planning regulations compared to other aspects. Besides the benchmark, we present the largest-ever supervised fine-tuning (SFT) dataset, UrbanPlanText, comprising over 30,000 instruction pairs sourced from urban planning exams and textbooks. Our findings demonstrate that fine-tuned models exhibit enhanced performance in memorization tests and comprehension of urban planning knowledge, while there exists significant room for improvement, particularly in tasks requiring domain-specific terminology and reasoning. By making our benchmark, dataset, and associated evaluation and fine-tuning toolsets publicly available at https://github.com/tsinghua-fib-lab/PlanBench, we aim to catalyze the integration of LLMs into practical urban planning, fostering a symbiotic collaboration between human expertise and machine intelligence.

城市规划大模型评测数据集LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。