让大模型搞定脑部MRI的多种临床任务,效果比专用模型还强。
Visual Instruction-Finetuned Language Model for Versatile Brain MR Image Tasks
- 用重用图像编码器特征缓解图像分块带来的信息损失
- 用大模型生成指令数据扩充稀缺的医图配对样本
- 在5个数据集上同时完成报告生成、问答、分割和图像重建
大语言模型在语言推理和视觉语言任务中表现卓越,通过将图像标记融入Transformer架构,已实现直接的视觉输入输出,推动研究从图像描述发展到文本生成图像。然而,简单的文本生成图像在临床应用中价值有限。在医学影像领域,如病灶定位的图像分割或缺失序列重建的图像转换等任务更具临床意义。尽管如此,如何将这些多样且关键的临床任务整合进一个通用的语言模型仍属空白。本文提出LLaBIT(用于脑部图像转换的大语言模型),将大模型的视觉推理能力扩展至脑部MRI中的多个临床相关任务。为缓解图像分块导致的空间信息丢失,我们引入机制复用图像编码器的特征图,减少数据退化;同时利用大模型按严格预设指令生成文本数据,以补充脑部MRI中有限的图像-文本配对数据。我们在五个脑部MRI数据集上,针对报告生成、视觉问答、图像分割和图像转换四项任务进行了全面评估。结果表明,该模型在所有任务中均表现优异,且在与专门优化的模型直接对比时也取得更优成绩,充分证明其高效性与多功能性。
原文摘要 · Abstract (English)
LLMs have demonstrated remarkable capabilities in linguistic reasoning and are increasingly adept at vision-language tasks. The integration of image tokens into transformers has enabled direct visual input and output, advancing research from image-to-text descriptions to text-to-image generation. However, simple text-to-image generation holds limited clinical utility. In medical imaging, tasks such as image segmentation for localizing pathologies or image translation for reconstructing missing sequences have much greater clinical importance. Despite this, integrating these diverse, clinically relevant tasks within a single, versatile language model remains unexplored. Our method, LLaBIT (Large Language Model for Brain Image Translation), extends the visual reasoning of LLMs to these clinically meaningful tasks in the brain MRI domain. To mitigate the spatial information loss inherent in image tokenization, we incorporate a mechanism to reuse feature maps from the image encoder, minimizing data degradation. We also generate text data using LLMs with strict predefined instructions to augment limited image-text paired data in brain MRI. We comprehensively evaluated our method on five brain MRI datasets across four distinct tasks: report generation, visual question answering, image segmentation, and image translation. Our model not only demonstrated superior performance across all tasks but also outperformed specialized, task-specific models in direct comparisons, highlighting its efficacy and versatility
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。