让医学影像模型真正理解解剖结构,提升胸部X光诊断准确率
AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation
- 分两阶段识别解剖结构并提取特征,再用大模型完成多任务推理
- 在解剖定位、短语定位等任务上性能提升超25%
- 适合需要精准解剖理解的临床辅助诊断场景
多模态医学大模型在胸部X光解读中取得显著进展,但在空间推理和解剖理解方面仍面临挑战。现有定位技术虽提升整体性能,但常无法建立真实的解剖对应关系,导致医学领域理解错误。为此,我们提出AnatomiX,一种面向解剖结构引导的胸部X光解读多任务多模态大模型。受放射科工作流程启发,AnatomiX采用两阶段方法:首先识别解剖结构并提取特征,随后利用大语言模型完成短语定位、报告生成、视觉问答和图像理解等下游任务。在多个基准上的广泛实验表明,AnatomiX在解剖推理方面表现优异,在解剖定位、短语定位、基于图像的诊断和基于图像的描述任务上相比现有方法性能提升超过25%。代码与预训练模型可在https://aneesurhashmi.github.io/anatomix获取。
原文摘要 · Abstract (English)
Multimodal medical large language models have shown substantial progress in chest X-ray interpretation but continue to face challenges in spatial reasoning and anatomical understanding. Although existing grounding techniques improve overall performance, they often fail to establish a true anatomical correspondence, resulting in incorrect anatomical understanding in the medical domain. To address this gap, we introduce AnatomiX, a multitask multimodal large language model for anatomically grounded chest X-ray interpretation. Inspired by the radiological workflow, AnatomiX adopts a two stage approach: first, it identifies anatomical structures and extracts their features, and then leverages a large language model to perform diverse downstream tasks such as phrase grounding, report generation, visual question answering, and image understanding. Extensive experiments across multiple benchmarks demonstrate that AnatomiX achieves superior anatomical reasoning and delivers over 25% improvement in performance on anatomy grounding, phrase grounding, grounded diagnosis and grounded captioning tasks compared to existing approaches. Code and pretrained model are available at https://aneesurhashmi.github.io/anatomix
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。