用AI双向生成胸片与报告,提升医学影像理解精度。
A Generative Framework for Bidirectional Image-Report Understanding in Chest Radiography
- 分阶段自适应调优,结合临床权重分词提升图文对齐
- 在MIMIC-CXR和印第安纳大学数据集上均达顶尖表现
- 适合医学AI研究者及临床辅助诊断系统开发者
大型语言模型(LLMs)在多模态任务中展现出巨大潜力,但将其应用于胸部X光(CXR)时面临视觉-文本精确对齐和关键诊断细节保留的挑战。本文提出多阶段自适应视觉-语言调优框架MAViLT,通过临床梯度加权标记化和分层微调策略,实现准确的放射科报告生成、基于文本的逼真胸片合成以及视觉临床问题回答。我们在MIMIC-CXR和印第安纳大学胸片数据集上评估了MAViLT,所有任务均达到当前最优性能。人工评估进一步验证了其临床相关性与实用性,证明该框架在真实医疗场景中的可行性,为利用LLMs实现多模态医学影像理解提供了有效路径。
原文摘要 · Abstract (English)
The rapid advancements in large language models (LLMs) have unlocked their potential for multimodal tasks, where text and visual data are processed jointly. However, applying LLMs to medical imaging, particularly for chest X-rays (CXR), poses significant challenges due to the need for precise visual-textual alignment and the preservation of critical diagnostic details. In this paper, we propose Multi-Stage Adaptive Vision-Language Tuning (MAViLT), a novel framework designed to enhance multimodal reasoning and generation for CXR understanding. MAViLT incorporates a clinical gradient-weighted tokenization process and a hierarchical fine-tuning strategy, enabling it to generate accurate radiology reports, synthesize realistic CXRs from text, and answer vision-based clinical questions. We evaluate MAViLT on two benchmark datasets, MIMIC-CXR and Indiana University CXR, achieving state-of-the-art results across all tasks. Human evaluations further validate the clinical relevance and utility of MAViLT, making it a robust tool for real-world medical applications. This work demonstrates the feasibility of leveraging LLMs for multimodal medical imaging while addressing key challenges in vision-language integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。