arXiv:2510.12400cs.CV2025-10综述被引 3

用视觉语言模型让机器像市民一样看懂城市设施状态。

Towards General Urban Monitoring with Vision-Language Models: A Review, Evaluation, and a Research Agenda

  • 用视觉语言模型实现无需训练的都市设施视觉判断。
  • 在32项研究中验证了零样本识别垃圾箱、路牌等的可行性。
  • 适合城市治理、智能交通与公共安全研究者参考。

城市公共基础设施(如垃圾箱、路标、植被、人行道和建筑工地)的监测因对象多样、环境复杂而面临挑战。当前主流方法依赖物联网传感器与人工巡检,成本高、难扩展,且常与市民直观观察不一致。这引发关键问题:机器能否像市民一样“看见”并判断城市设施状况?视觉语言模型(VLMs)融合视觉理解与自然语言推理,近期在处理复杂视觉信息方面表现突出,成为解决此问题的潜在技术。本系统性综述基于PRISMA方法,分析了2021至2025年间发表的32篇同行评审论文,回答四个核心问题:(1) VLMs已成功应用于哪些城市监测任务?(2) 哪些架构与框架表现更优?(3) 当前支持该领域的数据集与资源有哪些?(4) 如何评估基于VLM的应用,性能达到何种水平?

原文摘要 · Abstract (English)

Urban monitoring of public infrastructure (such as waste bins, road signs, vegetation, sidewalks, and construction sites) poses significant challenges due to the diversity of objects, environments, and contextual conditions involved. Current state-of-the-art approaches typically rely on a combination of IoT sensors and manual inspections, which are costly, difficult to scale, and often misaligned with citizens' perception formed through direct visual observation. This raises a critical question: Can machines now "see" like citizens and infer informed opinions about the condition of urban infrastructure? Vision-Language Models (VLMs), which integrate visual understanding with natural language reasoning, have recently demonstrated impressive capabilities in processing complex visual information, turning them into a promising technology to address this challenge. This systematic review investigates the role of VLMs in urban monitoring, with particular emphasis on zero-shot applications. Following the PRISMA methodology, we analyzed 32 peer-reviewed studies published between 2021 and 2025 to address four core research questions: (1) What urban monitoring tasks have been effectively addressed using VLMs? (2) Which VLM architectures and frameworks are most commonly used and demonstrate superior performance? (3) What datasets and resources support this emerging field? (4) How are VLM-based applications evaluated, and what performance levels have been reported?

城市监测视觉语言模型零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。