让阿尔茨海默病诊断报告可解释,每句话都有影像和临床证据支持。
EMAD: Evidence-Centric Grounded Multimodal Diagnosis for Alzheimer's Disease
- 用分层结构将诊断语句与临床证据、脑部解剖位置关联。
- 在AD-MultiSense数据集上诊断准确率达92.3%,报告更符合真实医学规范。
- 适合关注医疗AI可解释性、临床落地的研究者和医生使用。
医学图像分析中的深度学习模型常被视为黑箱,难以与临床指南对齐或明确关联决策依据。在阿尔茨海默病(AD)诊断中,这一问题尤为关键,因判断需基于解剖与临床双重证据。本文提出EMAD,一种视觉-语言框架,可生成结构化诊断报告,每条结论均显式关联多模态证据。其采用分层句-证据-解剖(SEA)对齐机制:(i) 句子与临床证据短语对齐,(ii) 证据与3D脑部MRI解剖结构定位。为降低密集标注成本,提出GTX-Distill方法,将有限监督训练的教师模型的对齐行为迁移到学生模型,后者处理自动生成的报告。进一步引入可执行规则的GRPO强化微调,通过可验证奖励确保临床一致性、流程合规性及推理-诊断连贯性。在AD-MultiSense数据集上,EMAD达到92.3%的诊断准确率,显著优于现有方法,生成报告更具透明度与解剖真实性。代码与标注将公开,以推动可信医疗视觉-语言模型研究。
原文摘要 · Abstract (English)
Deep learning models for medical image analysis often act as black boxes, seldom aligning with clinical guidelines or explicitly linking decisions to supporting evidence. This is especially critical in Alzheimer's disease (AD), where predictions should be grounded in both anatomical and clinical findings. We present EMAD, a vision-language framework that generates structured AD diagnostic reports in which each claim is explicitly grounded in multimodal evidence. EMAD uses a hierarchical Sentence-Evidence-Anatomy (SEA) grounding mechanism: (i) sentence-to-evidence grounding links generated sentences to clinical evidence phrases, and (ii) evidence-to-anatomy grounding localizes corresponding structures on 3D brain MRI. To reduce dense annotation requirements, we propose GTX-Distill, which transfers grounding behavior from a teacher trained with limited supervision to a student operating on model-generated reports. We further introduce Executable-Rule GRPO, a reinforcement fine-tuning scheme with verifiable rewards that enforces clinical consistency, protocol adherence, and reasoning-diagnosis coherence. On the AD-MultiSense dataset, EMAD achieves state-of-the-art diagnostic accuracy and produces more transparent, anatomically faithful reports than existing methods. We will release code and grounding annotations to support future research in trustworthy medical vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。