用Vision Transformer检测胶囊内镜中的17种罕见病,提升诊断效率。
RARE disease detection from Capsule Endoscopic Videos based on Vision Transformers
- 基于ViT的Transformer模型,处理224×224分辨率视频帧。
- 在3个测试视频上[email protected]为0.0205,[email protected]为0.0196。
- 适用于医学影像中多标签罕见病自动识别,辅助医生诊断。
本工作针对胶囊内镜视频(CEV)的多标签分类任务参与胃肠道竞赛。采用基于Transformer的深度学习网络进行微调,基础模型为Google Vision Transformer(ViT),批次大小为16,输入分辨率为224×224。共需分类17个标签:口、食管、胃、小肠、结肠、贲门、幽门、回盲瓣、活动性出血、血管扩张、血液、糜烂、充血、血红蛋白、淋巴管扩张、息肉和溃疡。在3个测试视频上,总体[email protected]为0.0205,[email protected]为0.0196。
原文摘要 · Abstract (English)
This work is corresponding to the Gastro Competition for multi-label classification from capsule endoscopic videos (CEV). Deep learning network based on Transformers are fined-tune for this task. The based online mode is Google Vision Transformer (ViT) batch16 with 224 x 224 resolutions. In total, 17 labels are classified, which are mouth, esophagus, stomach, small intestine, colon, z-line, pylorus, ileocecal valve, active bleeding, angiectasia, blood, erosion, erythema, hematin, lymphangioectasis, polyp, and ulcer. For test dataset of three videos, the overall mAP @0.5 is 0.0205 whereas the overall mAP @0.95 is 0.0196.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。