arXiv:2409.13321cs.LGcs.AI2024-09被引 7

轻量级医学影像报告生成模型,兼顾隐私与效率

SLaVA-CXR: Small Language and Vision Assistant for Chest X-ray Report Automation

  • 采用模拟放射科医生认知发展的三阶段训练法
  • 2.7B参数模型推理速度是先进模型的6倍
  • 自研数据合成方法保障隐私合规且提升多样性

受大语言模型成功的启发,医疗领域对LLMs辅助临床的研究日益增多。然而,医院使用闭源商业LLMs存在隐私风险,而开发开源公共LLMs又需大量计算资源,尤其在资源受限地区和低收入国家难以实现。本文提出开源小型语言视觉助手SLaVA-CXR,用于胸部X光报告自动化。为高效训练小型模型,我们提出Re$^3$Training方法,模拟放射科医生的认知发展过程,以识别、推理、报告三阶段优化模型;同时引入数据合成方法RADEX,可生成高质量且多样化的训练语料,符合隐私法规要求。大量实验表明,基于2.7B参数骨干网络的SLaVA-CXR不仅性能更优,且推理效率达同类先进模型的6倍。

原文摘要 · Abstract (English)

Inspired by the success of large language models (LLMs), there is growing research interest in developing LLMs in the medical domain to assist clinicians. However, for hospitals, using closed-source commercial LLMs involves privacy issues, and developing open-source public LLMs requires large-scale computational resources, which are usually limited, especially in resource-efficient regions and low-income countries. We propose an open-source Small Language and Vision Assistant (SLaVA-CXR) that can be used for Chest X-Ray report automation. To efficiently train a small assistant, we first propose the Re$^3$Training method, which simulates the cognitive development of radiologists and optimizes the model in the Recognition, Reasoning, and Reporting training manner. Then, we introduce a data synthesis method, RADEX, which can generate a high-quality and diverse training corpus with privacy regulation compliance. The extensive experiments show that our SLaVA-CXR built on a 2.7B backbone not only outperforms but also achieves 6 times faster inference efficiency than previous state-of-the-art larger models.

医学影像小模型报告生成隐私合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。