用视觉问答技术分析胃肠镜图像,提升医生诊断效率
Querying GI Endoscopy Images: A VQA Approach
- 将Florence2模型适配到胃肠镜医学图像的视觉问答任务
- 在专业医疗数据集上实现可验证的问答准确率提升
- 适合医学AI研发者和内镜诊断辅助系统开发者
视觉问答(VQA)融合自然语言处理与图像理解,可帮助临床医生更准确高效地诊断胃肠疾病。尽管当前多模态大模型在通用领域表现优异,但在医学影像等专业领域性能显著下降。本研究针对ImageCLEFmed-MEDVQA-GI 2025子任务1,探索将Florence2模型应用于胃肠镜图像的医学视觉问答任务,并采用ROUGE、BLEU和METEOR等标准指标评估模型性能。
原文摘要 · Abstract (English)
VQA (Visual Question Answering) combines Natural Language Processing (NLP) with image understanding to answer questions about a given image. It has enormous potential for the development of medical diagnostic AI systems. Such a system can help clinicians diagnose gastro-intestinal (GI) diseases accurately and efficiently. Although many of the multimodal LLMs available today have excellent VQA capabilities in the general domain, they perform very poorly for VQA tasks in specialized domains such as medical imaging. This study is a submission for ImageCLEFmed-MEDVQA-GI 2025 subtask 1 that explores the adaptation of the Florence2 model to answer medical visual questions on GI endoscopy images. We also evaluate the model performance using standard metrics like ROUGE, BLEU and METEOR
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。