arXiv:2608.12677cs.AIcs.CV2026-08

用视觉语言模型分析蚊子飞行视频,精准识别登革热感染状态

The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis

  • 结合YOLO与CLIP,将图像与生物语义文本对齐进行分类
  • 帧级准确率98.54%,灵敏度99.91%,视频级实现全正确识别
  • 适合生物行为分析、传染病早期检测的研究者参考

从视频数据中检测蚊子感染相关的运动变化极具挑战,因其体型小、移动迅速且不规则,易受背景、光照和阴影影响,导致特征提取困难。本研究提出一种基于YOLO与对比语言-图像预训练(CLIP)的跨模态框架,用于分类未感染与登革热病毒2型(DENV2)感染蚊子的飞行帧。首先利用YOLO分离蚊子区域,随后在共享嵌入空间中将视频帧的视觉特征与具有生物学意义的文本提示对齐。通过监督双向对比学习进行微调,并基于帧级图文相似性进行分类。结果表明,该方法在帧级达到98.54%准确率和99.91%灵敏度;经时间聚合后,模型实现完整的视频级性能。消融实验显示,微调和CLIP表征对本领域至关重要,而文本分支主要提供语义对齐而非精度提升。研究证实,视觉语言模型可为从视频数据中分析感染相关生物行为提供有效框架。

原文摘要 · Abstract (English)

Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and irregularly, and are affected by environmental factors such as background, lighting, and shadows, which can make reliable feature extraction difficult. In this study, a YOLO- and Contrastive Language-Image Pre-training (CLIP)-based vision-language framework is proposed to classify mosquito flight frames of uninfected and Dengue virus serotype 2 (DENV2)-infected mosquitoes. First, YOLO is used to isolate mosquito regions from the background. Then, visual features extracted from video frames are aligned with biologically meaningful textual prompts in a shared embedding space. The multimodal model was fine-tuned using supervised bidirectional contrastive learning and evaluated through frame-level image-text similarity-based classification. The results show that the proposed method achieved 98.54% accuracy and 99.91% sensitivity at the frame level. After temporal aggregation of frame-level information, the model achieved complete video-level performance. The ablation results showed that fine-tuning and CLIP-based representations were essential for this domain, while the textual branch provided semantic image-text alignment rather than an accuracy advantage over the vision-only model. These findings suggest that vision-language models can provide a useful framework for analyzing infection-related biological behaviors from video data.

视频分析登革热多模态生物行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。