arXiv:2410.19944cs.CV2024-10被引 4

用多模态模型自动识别胶囊胃镜中的10类病变,提升诊断效率。

A Multimodal Approach For Endoscopic VCE Image Classification Using BiomedCLIP-PubMedBERT

  • 融合视觉变换器与PubMedBERT,实现图像与文本联合建模。
  • 在10类病变分类上达到高准确率,关键指标表现优异。
  • 适合医疗AI研究者与内镜诊断辅助系统开发者参考。

本文提出一种基于BiomedCLIP-PubMedBERT的多模态方法,用于视频胶囊内镜(VCE)图像的异常分类,旨在提升消化道医疗诊断效率。通过将PubMedBERT语言模型与视觉变换器(ViT)结合,对内镜图像进行处理,实现对10类特定病变——包括血管扩张、出血、糜烂、充血、异物、淋巴管扩张、息肉、溃疡、寄生虫和正常——的分类。工作流程包含图像预处理,并对BiomedCLIP模型进行微调,生成高质量的视觉与文本嵌入,通过相似性评分对齐后完成分类。性能评估显示,该模型在分类准确率、召回率及F1分数等指标上均表现出色,展现出在临床诊断中实际应用的潜力。

原文摘要 · Abstract (English)

This Paper presents an advanced approach for fine-tuning BiomedCLIP PubMedBERT, a multimodal model, to classify abnormalities in Video Capsule Endoscopy (VCE) frames, aiming to enhance diagnostic efficiency in gastrointestinal healthcare. By integrating the PubMedBERT language model with a Vision Transformer (ViT) to process endoscopic images, our method categorizes images into ten specific classes: angioectasia, bleeding, erosion, erythema, foreign body, lymphangiectasia, polyp, ulcer, worms, and normal. Our workflow incorporates image preprocessing and fine-tunes the BiomedCLIP model to generate high-quality embeddings for both visual and textual inputs, aligning them through similarity scoring for classification. Performance metrics, including classification, accuracy, recall, and F1 score, indicate the models strong ability to accurately identify abnormalities in endoscopic frames, showing promise for practical use in clinical diagnostics.

多模态医学影像胶囊内镜分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。