打造高效医疗多模态模型,解决医学问答与诊断中的幻觉与效率问题。
InfiMed-Foundation: Pioneering Advanced Multimodal Medical Models with Compute-Efficient Pre-Training and Multi-Stage Fine-Tuning
- 用五维质量评估筛选高质量医研数据,提升训练数据可信度。
- 通过分阶段微调与多分辨率图像训练,实现40亿参数模型在医疗任务上的领先表现。
- 适合医疗AI研发者、临床辅助系统开发者使用,尤其关注诊断准确率提升。
多模态大语言模型在多个领域展现潜力,但在医疗应用中受限于专业知识不足、知识蒸馏效果差及大规模数据持续预训练的高算力需求。为此,我们提出InfiMed-Foundation-1.7B与InfiMed-Foundation-4B两款专用于医疗的多模态大模型。结合高质量通用与医学多模态数据,设计五维质量评估框架以筛选优质医疗数据集;采用低到高分辨率图像与多模态序列打包策略,显著提升训练效率,支持海量医学数据整合;并引入三阶段监督微调流程,有效提取复杂医疗任务所需知识。在MedEvalKit评测框架下,InfiMed-Foundation-1.7B超越Qwen2.5VL-3B,InfiMed-Foundation-4B优于HuatuoGPT-V-7B和MedGemma-27B-IT,尤其在医疗视觉问答与诊断任务中表现卓越。该工作为医疗AI提供更可靠、高效的解决方案。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have shown remarkable potential in various domains, yet their application in the medical field is hindered by several challenges. General-purpose MLLMs often lack the specialized knowledge required for medical tasks, leading to uncertain or hallucinatory responses. Knowledge distillation from advanced models struggles to capture domain-specific expertise in radiology and pharmacology. Additionally, the computational cost of continual pretraining with large-scale medical data poses significant efficiency challenges. To address these issues, we propose InfiMed-Foundation-1.7B and InfiMed-Foundation-4B, two medical-specific MLLMs designed to deliver state-of-the-art performance in medical applications. We combined high-quality general-purpose and medical multimodal data and proposed a novel five-dimensional quality assessment framework to curate high-quality multimodal medical datasets. We employ low-to-high image resolution and multimodal sequence packing to enhance training efficiency, enabling the integration of extensive medical data. Furthermore, a three-stage supervised fine-tuning process ensures effective knowledge extraction for complex medical tasks. Evaluated on the MedEvalKit framework, InfiMed-Foundation-1.7B outperforms Qwen2.5VL-3B, while InfiMed-Foundation-4B surpasses HuatuoGPT-V-7B and MedGemma-27B-IT, demonstrating superior performance in medical visual question answering and diagnostic tasks. By addressing key challenges in data quality, training efficiency, and domain-specific knowledge extraction, our work paves the way for more reliable and effective AI-driven solutions in healthcare. InfiMed-Foundation-4B model is available at \href{https://huggingface.co/InfiX-ai/InfiMed-Foundation-4B}{InfiMed-Foundation-4B}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。