arXiv:2511.18640cs.CVcs.AI2025-11被引 5

用医院日常影像数据训练出能懂脑部解剖与病变的通用医学模型

Health system learning achieves generalist neuroimaging models

  • 在524万份临床影像上训练,通过健康系统学习获取医学知识
  • 在诊断与报告生成任务中表现超越前沿模型,准确率显著提升
  • 适合医疗AI研究者与临床医生,助力安全可靠的智能辅助决策

前沿人工智能模型如OpenAI的GPT-5和Meta的DINOv3虽发展迅速,但受限于无法访问私有临床数据。神经影像因包含可识别面部特征,在公开数据中严重缺失,制约了其在临床中的应用。本文发现这些模型在神经影像任务中表现不佳,而通过健康系统学习——即直接利用医疗机构日常诊疗中生成的未标注数据,可构建高性能、通用型神经影像模型。我们提出NeuroVFM,一个基于524万例临床MRI与CT影像,采用可扩展的体积分组嵌入预测架构训练的视觉基础模型。该模型学习到完整的脑部解剖与病理表征,在多项临床任务中达到当前最优表现,包括放射学诊断与报告生成。模型展现出涌现的神经解剖理解能力与可解释的视觉定位能力。结合开源语言模型进行轻量级视觉指令微调后,其生成的放射科报告在准确性、临床分诊和专家偏好上均优于前沿模型。通过临床语境下的视觉理解,有效减少幻觉与关键错误,提供更安全的临床决策支持。本研究确立了健康系统学习作为通用医学AI的新范式,并提供了可扩展的临床基础模型框架。

原文摘要 · Abstract (English)

Frontier artificial intelligence (AI) models, such as OpenAI's GPT-5 and Meta's DINOv3, have advanced rapidly through training on internet-scale public data, yet such systems lack access to private clinical data. Neuroimaging, in particular, is underrepresented in the public domain due to identifiable facial features within MRI and CT scans, fundamentally restricting model performance in clinical medicine. Here, we show that frontier models underperform on neuroimaging tasks and that learning directly from uncurated data generated during routine clinical care at health systems, a paradigm we call health system learning, yields high-performance, generalist neuroimaging models. We introduce NeuroVFM, a visual foundation model trained on 5.24 million clinical MRI and CT volumes using a scalable volumetric joint-embedding predictive architecture. NeuroVFM learns comprehensive representations of brain anatomy and pathology, achieving state-of-the-art performance across multiple clinical tasks, including radiologic diagnosis and report generation. The model exhibits emergent neuroanatomic understanding and interpretable visual grounding of diagnostic findings. When paired with open-source language models through lightweight visual instruction tuning, NeuroVFM generates radiology reports that surpass frontier models in accuracy, clinical triage, and expert preference. Through clinically grounded visual understanding, NeuroVFM reduces hallucinated findings and critical errors, offering safer clinical decision support. These results establish health system learning as a paradigm for building generalist medical AI and provide a scalable framework for clinical foundation models.

医学影像健康系统学习视觉基础模型临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。