LoV3D通过区域体积分析实现脑部MRI纵向诊断,减少AI误判。
LoV3D: Grounding Cognitive Prognosis Reasoning in Longitudinal 3D Brain MRI via Regional Volume Assessments
- 分步处理脑MRI,结合解剖评估与时间对比。
- 三分类准确率93.7%,较基线提升34.8%。
- 无需人工标注,可跨数据集通用,适合临床研究。
纵向脑部MRI对阿尔茨海默病等神经退行性疾病的进展评估至关重要。当前深度学习工具存在割裂问题:分类器仅输出标签,体积分析产生无解释结果,视觉语言模型可能生成看似合理实则错误的结论。我们提出LoV3D,一种用于训练3D视觉语言模型的流程,可读取纵向T1加权脑MRI,进行区域级解剖评估,对比前序扫描,并输出三类诊断(认知正常、轻度认知障碍、痴呆)及合成摘要。该流程通过强制标签一致性、纵向连贯性和生物学合理性来约束最终诊断,降低幻觉风险。训练引入临床加权验证器,基于标准化体积指标自动评分候选输出,驱动直接偏好优化,全程无需人工标注。在保留测试集ADNI上(479次扫描,258名受试者),LoV3D达到93.7%三分类准确率(较无约束基线提升34.8%),两分类准确率97.2%(较最先进方法提升4%),区域解剖分类准确率82.6%(较VLM基线提升33.1%)。零样本迁移在MIRIAD上达95.4%准确率(100%痴呆召回率),AIBL上三分类准确率82.9%,证明其在不同站点、设备和人群间具有高泛化能力。代码已公开于https://github.com/Anonymous-TEVC/LoV-3D。
原文摘要 · Abstract (English)
Longitudinal brain MRI is essential for characterizing the progression of neurological diseases such as Alzheimer's disease assessment. However, current deep-learning tools fragment this process: classifiers reduce a scan to a label, volumetric pipelines produce uninterpreted measurements, and vision-language models (VLMs) may generate fluent but potentially hallucinated conclusions. We present LoV3D, a pipeline for training 3D vision-language models, which reads longitudinal T1-weighted brain MRI, produces a region-level anatomical assessment, conducts longitudinal comparison with the prior scan, and finally outputs a three-class diagnosis (Cognitively Normal, Mild Cognitive Impairment, or Dementia) along with a synthesized diagnostic summary. The stepped pipeline grounds the final diagnosis by enforcing label consistency, longitudinal coherence, and biological plausibility, thereby reducing the risks of hallucinations. The training process introduces a clinically-weighted Verifier that scores candidate outputs automatically against normative references derived from standardized volume metrics, driving Direct Preference Optimization without a single human annotation. On a subject-level held-out ADNI test set (479 scans, 258 subjects), LoV3D achieves 93.7% three-class diagnostic accuracy (+34.8% over the no-grounding baseline), 97.2% on two-class diagnosis accuracy (+4% over the SOTA) and 82.6% region-level anatomical classification accuracy (+33.1% over VLM baselines). Zero-shot transfer yields 95.4% on MIRIAD (100% Dementia recall) and 82.9% three-class accuracy on AIBL, confirming high generalizability across sites, scanners, and populations. Code is available at https://github.com/Anonymous-TEVC/LoV-3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。