HealthGPT统一医疗视觉理解与生成,用新方法提升模型性能。
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
- 通过异构低秩适配技术逐步融合多源医学知识
- 在VL-Health数据集上实现医疗视觉任务的卓越表现
- 适合医疗AI研发、临床辅助系统开发者使用
我们提出HealthGPT,一个强大的医学大视觉语言模型(Med-LVLM),在统一的自回归框架中整合医疗视觉理解与生成能力。基于渐进式适应异构理解与生成知识的构建理念,采用新颖的异构低秩适配(H-LoRA)技术,结合定制化的分层视觉感知方法和三阶段学习策略。为有效训练HealthGPT,我们构建了名为VL-Health的综合性医学领域专用理解与生成数据集。实验结果表明,HealthGPT在医疗视觉统一任务中表现出色且具备良好可扩展性。项目开源地址:https://github.com/DCDmllm/HealthGPT。
原文摘要 · Abstract (English)
We present HealthGPT, a powerful Medical Large Vision-Language Model (Med-LVLM) that integrates medical visual comprehension and generation capabilities within a unified autoregressive paradigm. Our bootstrapping philosophy is to progressively adapt heterogeneous comprehension and generation knowledge to pre-trained large language models (LLMs). This is achieved through a novel heterogeneous low-rank adaptation (H-LoRA) technique, which is complemented by a tailored hierarchical visual perception approach and a three-stage learning strategy. To effectively learn the HealthGPT, we devise a comprehensive medical domain-specific comprehension and generation dataset called VL-Health. Experimental results demonstrate exceptional performance and scalability of HealthGPT in medical visual unified tasks. Our project can be accessed at https://github.com/DCDmllm/HealthGPT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。