arXiv:2412.04954cs.CVcs.CL2024-12中稿 · ACL被引 8

用视觉语言模型生成胸部X光报告,准确率高。

Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation

  • 结合图像编码器与微调的LLM,理解胸部X光图像
  • 两阶段训练:先对齐图像特征,再微调生成报告
  • 适合医学影像生成、临床辅助诊断研究者

我们提出一种面向放射科的视觉语言模型,可基于胸部X光片生成放射科报告。基于先前研究中大型语言模型(LLMs)通过与预训练视觉编码器对齐可获得多模态能力的发现,我们验证了该方法在胸部X光图像上的适用性。该模型融合了图像编码器与基于Vicuna-7B架构的微调大语言模型,能够以较高准确性生成放射科报告的不同部分。训练采用两阶段策略:(i) 将胸部X光特征与大语言模型初步对齐;(ii) 随后进行放射科报告生成的微调。

原文摘要 · Abstract (English)

We introduce a radiology-focused visual language model designed to generate radiology reports from chest X-rays. Building on previous findings that large language models (LLMs) can acquire multimodal capabilities when aligned with pretrained vision encoders, we demonstrate similar potential with chest X-ray images. This integration enhances the ability of model to understand and describe chest X-ray images. Our model combines an image encoder with a fine-tuned LLM based on the Vicuna-7B architecture, enabling it to generate different sections of a radiology report with notable accuracy. The training process involves a two-stage approach: (i) initial alignment of chest X-ray features with the LLM (ii) followed by fine-tuning for radiology report generation.

医学影像视觉语言模型报告生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。