用3D视觉语言模型自动生成脑瘤MRI临床报告,准确率远超2D方法。
Brain3D: Brain Report Automation via Inflated Vision Transformers in 3D
- 将2D医疗编码器膨胀为原生3D结构,分三阶段对齐语言模型
- 在468例数据上临床病理F1达0.951,健康扫描特异性完美
- 专为神经放射学设计,擅长定位、侧向性和浸润模式分析
当前医学视觉-语言模型处理脑部MRI时仍依赖2D切片近似,破坏了神经放射学诊断所需的三维空间上下文。我们提出Brain3D,一种分阶段的视觉-语言框架,用于从3D脑瘤MRI自动生成放射科报告。该方法将预训练的2D医疗编码器膨胀为原生3D架构,并通过三个阶段逐步与因果语言模型对齐:对比性定位、监督投影器预热和基于LoRA的语言专业化。与通用3D医学VLM不同,Brain3D专为神经放射学优化,重点关注半球侧向性、肿瘤浸润模式和解剖定位。在468名受试者(BraTS病理病例及健康对照)上的评估显示,模型临床病理F1达到0.951,而强2D基线仅为0.413,同时对健康扫描保持完美特异性。分阶段对齐至关重要:对比性定位建立视觉-文本对应关系,投影器预热稳定条件输入,LoRA适配使输出从冗长描述转向结构化临床报告。
原文摘要 · Abstract (English)
Current medical vision-language models (VLMs) process volumetric brain MRI using 2D slice-based approximations, fragmenting the spatial context required for accurate neuroradiological interpretation. We developed \textbf{Brain3D}, a staged vision-language framework for automated radiology report generation from 3D brain tumor MRI. Our approach inflates a pretrained 2D medical encoder into a native 3D architecture and progressively aligns it with a causal language model through three stages: contrastive grounding, supervised projector warmup, and LoRA-based linguistic specialization. Unlike generalist 3D medical VLMs, \textbf{Brain3D} is tailored to neuroradiology, where hemispheric laterality, tumor infiltration patterns, and anatomical localization are critical. Evaluated on 468 subjects (BraTS pathological cases plus healthy controls), our model achieves a Clinical Pathology F1 of 0.951 versus 0.413 for a strong 2D baseline while maintaining perfect specificity on healthy scans. The staged alignment proves essential: contrastive grounding establishes visual-textual correspondence, projector warmup stabilizes conditioning, and LoRA adaptation shifts output from verbose captions to structured clinical reports\footnote{Our code is publicly available for transparency and reproducibility
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。