Hulu-Med统一处理医学文本、图像、视频等多模态数据,性能超越多数开源与闭源模型。
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
- 设计通用医学多模态模型,整合文本、2D/3D图像与视频理解于单一架构。
- 在30个医学基准上,27项优于现有开源模型,16项超越GPT-4o。
- 首次公开完整训练流程与模型参数,实现透明可复现的医疗多模态研究。
真实临床决策需融合异构数据,包括医学文本、2D图像、3D体数据和视频,但现有AI系统无法统一处理所有信号,限制其应用。本文提出Hulu-Med,一个透明的通用医学视觉-语言模型(VLM),可在单一架构中统一处理纯文本、2D/3D视觉-语言及视频理解任务。该模型基于1670万条精选样本训练,全部为公开或合成数据,覆盖12个主要解剖系统和14种医学成像模态。Hulu-Med采用医学感知的令牌压缩策略,对3D和视频输入可减少高达55%冗余视觉令牌,提升跨模态效率,支持7B–32B参数规模模型在约4000–40000 GPU小时内的训练。在30个公共域内与域外医学基准测试中——涵盖文本推理、视觉问答、报告生成、多语言对话、视频理解与罕见病诊断——Hulu-Med在27项中超越现有开源模型,在16项中优于专有系统GPT-4o。尽管是视觉-语言模型,其在纯文本健康基准HealthBench上的表现仍超过GPT-4o,并与GPT-o1持平。本工作首次向社区提供端到端透明、可复现且成本可控的全息医学多模态理解方案,包含数据清洗、训练流程与模型参数。代码与模型已开源:https://github.com/ZJUI-AI4H/Hulu-Med。
原文摘要 · Abstract (English)
Real-world clinical decision-making requires integrating heterogeneous data, including medical text, 2D images, 3D volumes, and videos, while existing AI systems fail to unify all these signals, limiting their utility. In this paper, we introduce Hulu-Med, a transparent, generalist medical Vision-Language Model (VLM) designed to unify language-only, 2D/3D vision-language, and video understanding within a single architecture. Hulu-Med is trained on a curated corpus of 16.7 million samples, comprising exclusively public or synthetic data, spanning 12 major anatomical systems and 14 medical imaging modalities. Hulu-Med employs a medical-aware token-reduction strategy that prunes redundant visual tokens, achieving up to a 55% reduction for 3D and video inputs, improving cross-modal efficiency, and enabling training at 7B-32B parameter scales in approximately 4,000-40,000 GPU hours. Across 30 public in-domain and out-of-domain medical benchmarks-covering text reasoning, visual question answering, report generation, multilingual dialogue, video understanding, and rare disease diagnosis-Hulu-Med surpasses existing open-source models on 27 of 30 benchmarks and outperforms proprietary systems such as GPT-4o on 16 benchmarks. Despite being a VLM, Hulu-Med outperforms GPT-4o and matches GPT-o1 on the text-only HealthBench. For the first time in the community, we provide a fully transparent, reproducible and cost-effective pipeline for holistic medical vision-language understanding by releasing our end-to-end data curation, training procedures, and model parameters. Code and models are available at https://github.com/ZJUI-AI4H/Hulu-Med.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。