arXiv:2512.14312cs.CVcs.AI2025-12

用视觉语言模型零样本识别中东污水厂,比传统方法更高效。

From YOLO to VLMs: Advancing Zero-Shot and Few-Shot Detection of Wastewater Treatment Plants Using Satellite Imagery in MENA Region

  • 用多模态大模型零样本/少样本识别卫星图像中的污水处理设施
  • Gemma-3在零样本下超越YOLOv8的真阳性率,达85%以上
  • 适合遥感、环境监测领域研究者快速部署无标注检测系统

中东和北非(MENA)地区对污水处理厂(WWTP)需求迫切,精准识别有助于可持续水资源管理。传统方法如YOLOv8需大量人工标注,而视觉语言模型(VLMs)可通过内在推理实现高效替代。本研究构建了零样本与少样本双路径对比框架,基于埃及、沙特阿拉伯和阿联酋的83,566张高分辨率卫星图像训练YOLOv8(约85%为正样本,15%为负样本)。评估模型包括LLaMA 3.2 Vision、Qwen 2.5 VL、DeepSeek-VL2、Gemma 3、Gemini及Pixtral 12B(Mistral),通过专家提示识别池体、曝气池等组件,并区分干扰项,输出含置信度与描述的JSON结果。数据集包含1,207个经验证的WWTP位置(阿联酋198个、沙特354个、埃及655个),对应数量的非WWTP点位,均为600m×600m Geo-TIFF图像(缩放级别18,坐标系EPSG:4326)。零样本测试显示多个VLM优于YOLOv8的真阳性率,其中Gemma-3表现最佳。结果表明,尤其在零样本条件下,VLM可有效替代YOLOv8,实现无需标注的大规模远程感知。

原文摘要 · Abstract (English)

In regions of the Middle East and North Africa (MENA), there is a high demand for wastewater treatment plants (WWTPs), crucial for sustainable water management. Precise identification of WWTPs from satellite images enables environmental monitoring. Traditional methods like YOLOv8 segmentation require extensive manual labeling. But studies indicate that vision-language models (VLMs) are an efficient alternative to achieving equivalent or superior results through inherent reasoning and annotation. This study presents a structured methodology for VLM comparison, divided into zero-shot and few-shot streams specifically to identify WWTPs. The YOLOv8 was trained on a governmental dataset of 83,566 high-resolution satellite images from Egypt, Saudi Arabia, and UAE: ~85% WWTPs (positives), 15% non-WWTPs (negatives). Evaluated VLMs include LLaMA 3.2 Vision, Qwen 2.5 VL, DeepSeek-VL2, Gemma 3, Gemini, and Pixtral 12B (Mistral), used to identify WWTP components such as circular/rectangular tanks, aeration basins and distinguish confounders via expert prompts producing JSON outputs with confidence and descriptions. The dataset comprises 1,207 validated WWTP locations (198 UAE, 354 KSA, 655 Egypt) and equal non-WWTP sites from field/AI data, as 600mx600m Geo-TIFF images (Zoom 18, EPSG:4326). Zero-shot evaluations on WWTP images showed several VLMs out-performing YOLOv8's true positive rate, with Gemma-3 highest. Results confirm that VLMs, particularly with zero-shot, can replace YOLOv8 for efficient, annotation-free WWTP classification, enabling scalable remote sensing.

视觉语言模型遥感检测零样本学习污水厂识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。