arXiv:2606.11678cs.CL2026-06

测试AI能否像城市规划师一样思考,发现它在复杂判断上表现反而优于基础记忆。

Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment

论文配图:Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment
图 1 · 摘自论文原文
  • 用4大知识维度+5层认知水平构建评估框架,系统评测25个大模型
  • 模型在高阶分析任务上表现好,但低阶事实记忆和综合判断能力差
  • 揭示AI在制度、时空语境理解上的四大认知局限,适合辅助而非替代决策

大语言模型(LLMs)的兴起引发关键问题:城市规划中的哪些专业知识可被AI复制,哪些仍需人类判断?尽管AI工具日益应用于规划实践,却缺乏系统性框架来检验其是否具备规划专业所依赖的情境敏感性、价值意识与制度素养。本文提出专用于城市规划领域的评估框架UPBench,基于布卢姆修订版认知分类法设计4×5矩阵,涵盖四大知识支柱与五级认知层次。通过自动化评分与专家评审,对25个主流大模型进行评估,发现认知能力呈现非单调曲线:模型在高阶分析任务上表现优于事实回忆与整合判断。这表明规划知识中常被视为低阶的部分,实则深度嵌入制度、管辖权与时间语境,导致模型难以泛化。我们总结出四类认知局限:监管幻觉、概念混淆、复杂难题瘫痪、实践智慧缺失。实践启示:建议对规划工作实行差异化分工——大模型可协助跨学科整合、文献综述、情景生成与初步政策分析,但在具体法规适配、规范冲突调解与情境敏感流程方面仍不可靠。机构应要求对AI辅助的法规分析进行人工验证,规划教育则需强化制度理解、规范判断与情境敏感性训练。

原文摘要 · Abstract (English)

Problem, Research Strategy, and Findings: The rise of large language models (LLMs) raises a key question for urban planning: which forms of professional planning knowledge can AI replicate, and which still require human judgment? Although AI tools are increasingly used in planning practice, there is still no systematic framework for testing whether they can reason with the contextual sensitivity, value awareness, and institutional literacy central to planning expertise. This paper introduces Urban Planning Bench (UPBench), a domain-specific evaluation framework that assesses LLM reasoning through a 4x5 matrix of four knowledge pillars and five cognitive levels adapted from Bloom's revised taxonomy. Evaluating 25 LLMs with automated scoring and expert review, we find a non-monotonic cognitive curve: models perform better on higher-order analytical tasks than on factual recall and integrative judgment. This suggests that planning knowledge often treated as lower-order is deeply shaped by institutional, jurisdictional, and temporal context, making it hard for LLMs to generalize. We summarize these limits as four epistemic diagnostics: regulatory hallucination, conceptual conflation, wickedness paralysis, and phronetic deficit. Takeaway for Practice: The findings support differential delegation in planning. LLMs can assist with cross-disciplinary synthesis, literature review, scenario generation, and preliminary policy analysis. However, they remain unreliable for jurisdiction-specific regulation, normative conflict resolution, and context-sensitive procedure. Agencies should require verification for AI-assisted regulatory analysis, while planning education should emphasize institutional literacy, normative judgment, and contextual sensitivity.

城市规划大模型评估认知局限AI辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。