arXiv:2605.04772cs.CV2026-05中稿 · the Workshop on Ap…被引 1

MIRAGE可检索生成可信医学图像,助力医学生互动学习。

MIRAGE: Retrieval and Generation of Multimodal Images and Texts for Medical Education

论文配图:MIRAGE: Retrieval and Generation of Multimodal Images and Texts for Medical Education
图 1 · 摘自论文原文
  • 将文本与图像映射到共享潜空间,支持语义化查询。
  • 基于公开预训练模型,支持图文检索与合成生成。
  • 无需编程技能,网页端即用,适合医学生快速上手。

获取多样且标注准确的医学图像及交互式学习工具,对提升医学生诊断能力和解剖结构理解至关重要。传统医学图谱因体积大且缺乏交互性难以使用,而网络图像搜索则可能提供错误或不完整的资料。为此,我们提出MIRAGE,一个用于医学教育的多模态文本与图像检索与生成系统。该系统通过微调的医学版CLIP(MedICaT-ROCO)将文本和图像映射至共享潜空间,实现语义相关的查询。系统支持用户输入提示词检索真实图像、通过医疗扩散模型Prompt2MedImage生成合成图像,并由大语言模型Dolly-v2-3b生成增强描述。还提供双模式搜索功能,便于对比不同疾病图像。系统完全依赖公开预训练模型,保证可复现性与可访问性。目标是为医学生提供免费、透明、易用的教学工具,尤其适用于无编程背景者。系统已部署于Kaggle平台,可通过网页界面实现交互式个性化视觉学习,无需本地计算资源或技术专长。

原文摘要 · Abstract (English)

Access to diverse, well-annotated medical images with interactive learning tools is fundamental for training practitioners in medicine and related fields to improve their diagnostic skills and understanding of anatomical structures. While medical atlases are valuable, they are often impractical due to their size and lack of interactivity, whereas online image search may provide mislabeled or incomplete material. To address this, we propose MIRAGE, a multimodal medical text and image retrieval and generation system that allows users to find and generate clinically relevant images from trustworthy sources by mapping both text and images to a shared latent space, enabling semantically meaningful queries. The system is based on a fine-tuned medical version of CLIP (MedICaT-ROCO), trained with the ROCO dataset, obtained from PubMed Central. MIRAGE allows users to give prompts to retrieve images, generate synthetic ones through a medical diffusion model (Prompt2MedImage) and receive enriched descriptions from a large language model (Dolly-v2-3b). It also supports a dual search option, enabling the visual comparison of different medical conditions. A key advantage of the system is that it relies entirely on publicly available pretrained models, ensuring reproducibility and accessibility. Our goal is to provide a free, transparent and easy-to-use didactic tool for medical students, especially those without programming skills. The system features an interface that enables interactive and personalized visual learning through medical image retrieval and generation. The system is accessible to medical students worldwide without requiring local computational resources or technical expertise, and is currently deployed on Kaggle: http://www-vpu.eps.uam.es/mirage

医学图像多模态生成模型教育工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。