用解剖学知识引导多步推理,提升医学影像模型的准确性和可解释性。
AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation

- 基于解剖本体构建跨模态区域级推理框架
- 在VQA和报告生成任务中显著优于现有方法
- 适合需要高可解释性的临床辅助诊断场景
胸部X光片(CXRs)是临床中最常见的影像检查。近年来,大型多模态模型(LMMs)在自动化胸部X光解读方面取得进展,提升了诊断准确率与效率。然而,当前医学多模态模型(MLMMs)仍面临两大挑战:(1) 区域级理解与交互不足;(2) 因单步推理导致准确率与可解释性受限。本文提出一种解剖本体引导推理(AOR)框架,以跨模态区域级信息为中心,支持多步推理,增强模型的交互性与可解释性。在专家医师指导下,我们构建了用于训练的大规模指令数据集AOR-Instruction。实验表明,AOR在视觉问答(VQA)和报告生成任务中均表现更优。
原文摘要 · Abstract (English)
Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Large Multimodal Models (LMMs) have enabled automated CXR interpretation, enhancing diagnostic accuracy and efficiency. However, despite their strong visual understanding, current Medical LMMs (MLMMs) still face two major challenges: (1) Insufficient region-level understanding and interaction, and (2) Limited accuracy and interpretability due to single-step reasoning. In this paper, we empower MLMMs with anatomy-centric reasoning capabilities to enhance their interactivity and explainability. Specifically, we first propose an Anatomical Ontology-Guided Reasoning (AOR) framework, which centers on cross-modal region-level information to facilitate multi-step reasoning. Next, under the guidance of expert physicians, we develop AOR-Instruction, a large instruction dataset for MLMMs training. Our experiments demonstrate AOR's superior performance in both VQA and report generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。