arXiv:2506.14629cs.CVcs.CL2025-06中稿 · CVPR被引 1

构建多模态数据集,实现蚊子滋生地的视觉检测、分割与文本解释。

VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites

  • 整合图像与文本,支持蚊子滋生地的检测、分割和自动解释。
  • 检测模型最高精度达0.929,分割模型mAP@50为0.798。
  • 适合公共卫生、智能监控与跨模态AI研究者使用。

蚊媒疾病对全球健康构成重大威胁,需通过早期发现和主动控制滋生地来预防疫情暴发。本文提出VisText-Mosquito,一个融合视觉与文本数据的多模态数据集,用于支持蚊子滋生地的自动化检测、分割与解释分析。数据集包含1,828张标注图像用于目标检测,142张用于水面分割,并为每张图像配套自然语言解释文本。YOLOv9s在检测任务中达到最高精度0.92926和mAP@50 0.92891;YOLOv11n-Seg在分割任务中取得精度0.91587和mAP@50 0.79795。在文本生成方面,测试了多种大视觉-语言模型(LVLM),微调后的Mosquito-LLaMA3-8B模型表现最优,最终损失0.0028,BLEU得分54.7,BERTScore 0.91,ROUGE-L 0.85。该数据集与模型框架体现了‘预防优于治疗’的理念,展示人工智能在主动防控蚊媒疾病中的潜力。数据集与代码已公开于GitHub:https://github.com/adnanul-islam-jisun/VisText-Mosquito。

原文摘要 · Abstract (English)

Mosquito-borne diseases pose a major global health risk, requiring early detection and proactive control of breeding sites to prevent outbreaks. In this paper, we present VisText-Mosquito, a multimodal dataset that integrates visual and textual data to support automated detection, segmentation, and explanation for mosquito breeding site analysis. The dataset includes 1,828 annotated images for object detection, 142 images for water surface segmentation, and natural language explanation texts linked to each image. The YOLOv9s model achieves the highest precision of 0.92926 and mAP@50 of 0.92891 for object detection, while YOLOv11n-Seg reaches a segmentation precision of 0.91587 and mAP@50 of 0.79795. For textual explanation generation, we tested a range of large vision-language models (LVLMs) in both zero-shot and few-shot settings. Our fine-tuned Mosquito-LLaMA3-8B model achieved the best results, with a final loss of 0.0028, a BLEU score of 54.7, BERTScore of 0.91, and ROUGE-L of 0.85. This dataset and model framework emphasize the theme "Prevention is Better than Cure", showcasing how AI-based detection can proactively address mosquito-borne disease risks. The dataset and implementation code are publicly available at GitHub: https://github.com/adnanul-islam-jisun/VisText-Mosquito

多模态蚊媒疾病目标检测文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。