arXiv:2502.19546cs.AIcs.CL2025-02被引 4

用医学论文训练神经外科视觉语言模型,效果接近大模型但更透明可信。

CNS-Obsidian: A Neurosurgical Vision-Language Model Built From Scientific Publications

  • 基于2.4万篇论文构建神经外科专用视觉语言模型
  • 在真实临床场景中诊断准确率低于GPT-4o但差距缩小
  • 为医学领域提供可解释、低成本的AI建模新范式

通用视觉语言模型虽能力强,但因训练数据未经筛选,难以用于高风险医疗决策。本文提出基于同行评审文献训练的神经外科视觉语言模型CNS-Obsidian。研究收集了来自《Neurosurgery Publications》期刊的23,984篇论文,包含78,853张图表及说明,利用GPT-4o和Claude Sonnet-3.5生成263,064个训练样本,涵盖指令微调、多选题与鉴别诊断三种格式。在此基础上,对340亿参数的LLaVA-Next模型进行微调。在2024年8月30日至11月30日于纽约大学朗格尼健康中心开展的盲法随机试验中,共评估70次神经外科会诊(占总会诊量959次的7.3%),由模型作为诊断辅助工具。主要评估指标为诊断帮助性与准确性。结果显示,合成问题上,CNS-Obsidian准确率为76.13%,接近GPT-4o的77.54%(p=0.235);但在人类生成的问题上,其准确率仅为46.81%,显著低于GPT-4o的65.70%(p<10^-15)。用户评分方面,CNS-Obsidian获得正面评价的比例为40.62%,低于GPT-4o的57.89%(p=0.230)。然而,两模型均在约60%情况下包含正确诊断(分别为59.38%与65.79%,p=0.626)。表明经过专业文献训练的领域特定模型,即使规模小、成本低,也能达到前沿模型水平,为科学社区构建可解释、可复现的专用人工智能提供了透明路径。

原文摘要 · Abstract (English)

General-purpose VLMs demonstrate impressive capabilities, but their opaque training on uncurated internet data poses critical limitations for high-stakes decision-making, such as in neurosurgery. We present CNS-Obsidian, a neurosurgical VLM trained on peer-reviewed literature, and demonstrate its clinical utility versus GPT-4o in a real-world setting. We compiled 23,984 articles from Neurosurgery Publications journals, yielding 78,853 figures and captions. Using GPT-4o and Claude Sonnet-3.5, we converted these into 263,064 training samples across three formats: instruction fine-tuning, multiple-choice questions, and differential diagnosis. We trained CNS-Obsidian, a fine-tune of the 34-billion parameter LLaVA-Next model. In a blinded, randomized trial at NYU Langone Health (Aug 30-Nov 30, 2024), neurosurgery consultations were assigned to either CNS-Obsidian or a HIPAA-compliant GPT-4o endpoint as diagnostic co-pilot after consultations. Primary outcomes were diagnostic helpfulness and accuracy, assessed via user ratings and presence of correct diagnosis within the VLM-provided differential. CNS-Obsidian matched GPT-4o on synthetic questions (76.13% vs 77.54%, p=0.235), but only achieved 46.81% accuracy on human-generated questions versus GPT-4o's 65.70% (p<10-15). In the randomized trial, 70 consultations were evaluated (32 CNS-Obsidian, 38 GPT-4o) from 959 total consults (7.3% utilization). CNS-Obsidian received positive ratings in 40.62% of cases versus 57.89% for GPT-4o (p=0.230). Both models included correct diagnosis in approximately 60% of cases (59.38% vs 65.79%, p=0.626). Domain-specific VLMs trained on curated scientific literature can approach frontier model performance despite being orders of magnitude smaller and less expensive to train. This establishes a transparent framework for scientific communities to build specialized AI models.

神经外科视觉语言模型医学AI可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。