用大模型精准识别道路设施状态,符合工程规范。
Unleashing the Capabilities of Large Vision-Language Models for Intelligent Perception of Roadside Infrastructure
- 用小样本微调+知识增强推理,让大模型懂工程规则。
- 检测mAP达58.9,属性识别准确率95.5%。
- 适合智慧城市建设者和交通系统开发者。
城市道路设施自动化感知对智慧城市建设至关重要,但通用模型难以捕捉细粒度属性与领域规范。尽管大视觉语言模型(VLMs)在开放世界识别中表现优异,却常因无法准确解读复杂设施状态而违背工程标准,导致实际应用可靠性差。为此,我们提出一种领域自适应框架,将VLMs转化为专业基础设施分析代理。方法结合数据高效微调与知识引导推理:先在Grounding DINO上采用开集微调实现多样资产的鲁棒定位,再基于LoRA对Qwen-VL进行深度语义属性推理适配;为减少幻觉并确保专业合规性,引入双模态检索增强生成(RAG)模块,在推理时动态获取权威行业标准与视觉范例。在新构建的城市道路场景数据集上评估,框架检测性能达58.9 mAP,属性识别准确率达95.5%,展现出可靠的道路设施智能监测能力。
原文摘要 · Abstract (English)
Automated perception of urban roadside infrastructure is crucial for smart city management, yet general-purpose models often struggle to capture the necessary fine-grained attributes and domain rules. While Large Vision Language Models (VLMs) excel at open-world recognition, they often struggle to accurately interpret complex facility states in compliance with engineering standards, leading to unreliable performance in real-world applications. To address this, we propose a domain-adapted framework that transforms VLMs into specialized agents for intelligent infrastructure analysis. Our approach integrates a data-efficient fine-tuning strategy with a knowledge-grounded reasoning mechanism. Specifically, we leverage open-vocabulary fine-tuning on Grounding DINO to robustly localize diverse assets with minimal supervision, followed by LoRA-based adaptation on Qwen-VL for deep semantic attribute reasoning. To mitigate hallucinations and enforce professional compliance, we introduce a dual-modality Retrieval-Augmented Generation (RAG) module that dynamically retrieves authoritative industry standards and visual exemplars during inference. Evaluated on a comprehensive new dataset of urban roadside scenes, our framework achieves a detection performance of 58.9 mAP and an attribute recognition accuracy of 95.5%, demonstrating a robust solution for intelligent infrastructure monitoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。