构建首个针对水体漂浮垃圾的多实例推理标注基准,推动轻量视觉语言模型精准识别。
WADE: A Reasoning-Annotated Benchmark for Multi-Instance Floating-Waste Grounding with Compact Vision-Language Models

- 构建含2167张图、13608个框的多实例垃圾数据集,附分类规则与推理链
- 轻量模型经微调后召回率提升至0.2339,幻觉率从0.6836降至0.0883
- 适合关注环境监测、小模型部署与多任务推理的研究者
内陆水道中的漂浮垃圾威胁水生生态系统,需在复杂多目标环境下及时监测。现有水体垃圾数据集地理覆盖有限、多实例标注稀疏,且监督信息仅限边界框与标签。因此,轻量视觉语言模型(VLMs)在联合定位、分类、计数和解释漂浮垃圾方面仍缺乏充分评估。本文提出WADE,一个包含2,167张来自孟加拉国农村的图像、13,608个边界框和十类垃圾的推理标注基准。每条标注关联类别级识别规则,涵盖视觉线索、易混淆项与判别特征。我们在零样本、两样本、推理引导及微调设置下,使用检测、计数与幻觉指标评估六种VLMs。为实现资源高效适配,我们采用QLoRA联合微调Qwen3-VL-2B模型,优化边界框、标签与推理链。微调后,召回率由0.0248提升至0.2339,F1值从0.0257增至0.2163,图像级幻觉率由0.6836降至0.0883。然而超过四分之三的实例仍未被检测,表明WADE对紧凑型VLMs具有显著挑战性。
原文摘要 · Abstract (English)
Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquatic-waste datasets provide limited geographic coverage, sparse multi-instance annotations, and little supervision beyond boxes and labels. Compact vision-language models (VLMs) therefore remain insufficiently evaluated for jointly localizing, classifying, counting, and explaining floating waste. We introduce WADE, a reasoning-annotated benchmark containing 2,167 images from rural Bangladesh, 13,608 bounding boxes, and ten waste categories. Each annotation is associated with class-level recognition rules covering visual cues, likely confusions, and discriminative features. We evaluate six VLMs under zero-shot, two-shot, reasoning-guided, and fine-tuned settings using detection, counting, and hallucination metrics. For resource-efficient adaptation, we jointly fine-tune Qwen3-VL-2B on boxes, labels, and reasoning chains using QLoRA. Fine-tuning increases recall from 0.0248 to 0.2339 and F1 from 0.0257 to 0.2163, while reducing image-level hallucination from 0.6836 to 0.0883. However, over three-quarters of instances remain undetected, establishing WADE as a challenging benchmark for dense floating-waste grounding with compact VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。