用视觉特征动态生成专属提示,提升医学影像报告生成效果。
Multimodal Large Language Models for Medical Report Generation via Customized Prompt Tuning
- 根据图像特征动态生成个性化提示,实现精准引导。
- 在IU X-ray和MIMIC-CXR数据集上达到当前最佳性能。
- 适合医学影像与自然语言生成交叉研究者使用。
医学影像报告生成在临床实践中仍具挑战性。尽管大语言模型(LLM)展现出巨大潜力,但其与医学影像数据的有效融合仍需深入探索。本文提出MRG-LLM,一种新型多模态大语言模型(MLLM),将冻结的LLM与可学习的视觉编码器结合,并引入动态提示定制机制。核心创新在于通过从视觉特征中提取的条件仿射变换,为每张医学图像生成实例化提示。我们提出了两种实现方式:按提示定制与按提示册定制,实现精准、针对性的报告生成。在IU X-ray和MIMIC-CXR数据集上的大量实验表明,MRG-LLM在医学报告生成任务中达到当前最优性能。代码将公开发布。
原文摘要 · Abstract (English)
Medical report generation from imaging data remains a challenging task in clinical practice. While large language models (LLMs) show great promise in addressing this challenge, their effective integration with medical imaging data still deserves in-depth exploration. In this paper, we present MRG-LLM, a novel multimodal large language model (MLLM) that combines a frozen LLM with a learnable visual encoder and introduces a dynamic prompt customization mechanism. Our key innovation lies in generating instance-specific prompts tailored to individual medical images through conditional affine transformations derived from visual features. We propose two implementations: prompt-wise and promptbook-wise customization, enabling precise and targeted report generation. Extensive experiments on IU X-ray and MIMIC-CXR datasets demonstrate that MRG-LLM achieves state-of-the-art performance in medical report generation. Our code will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。