arXiv:2409.12184cs.LGcs.AI2024-09被引 7

轻量化医疗多模态模型,可在低算力设备上高效运行

Democratizing MLLMs in Healthcare: TinyLLaVA-Med for Efficient Healthcare Diagnostics in Resource-Constrained Settings

  • 将TinyLLaVA适配为医疗专用模型TinyLLaVA-Med,通过指令微调优化
  • 在Jetson Xavier上仅需18.9W功耗和11.9GB内存,推理速度更快
  • 医疗问答任务准确率接近顶尖模型,适合偏远地区部署

在医疗领域部署多模态大模型(MLLMs)受限于其高计算需求与内存占用,尤其在资源有限的设备如Nvidia Jetson Xavier上更为显著。本文提出对通用多模态模型TinyLLaVA进行优化并改名为TinyLLaVA-Med,借鉴LLaVA-Med训练流程,在医疗数据集上进行指令微调与精调。该方法显著降低计算复杂度与功耗,使模型在18.9W功耗下运行,仅占用11.9GB内存,同时在闭式问题问答任务VQA-RAD上达到64.54%准确率,在SLAKE上达到70.70%。结果表明,TinyLLaVA-Med可在低算力硬件环境中实现可部署性,保持核心功能并接近当前最优模型性能。

原文摘要 · Abstract (English)

Deploying Multi-Modal Large Language Models (MLLMs) in healthcare is hindered by their high computational demands and significant memory requirements, which are particularly challenging for resource-constrained devices like the Nvidia Jetson Xavier. This problem is particularly evident in remote medical settings where advanced diagnostics are needed but resources are limited. In this paper, we introduce an optimization method for the general-purpose MLLM, TinyLLaVA, which we have adapted and renamed TinyLLaVA-Med. This adaptation involves instruction-tuning and fine-tuning TinyLLaVA on a medical dataset by drawing inspiration from the LLaVA-Med training pipeline. Our approach successfully minimizes computational complexity and power consumption, with TinyLLaVA-Med operating at 18.9W and using 11.9GB of memory, while achieving accuracies of 64.54% on VQA-RAD and 70.70% on SLAKE for closed-ended questions. Therefore, TinyLLaVA-Med achieves deployment viability in hardware-constrained environments with low computational resources, maintaining essential functionalities and delivering accuracies close to state-of-the-art models.

医疗AI轻量化模型多模态边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。