7B参数医学多模态模型压缩至4GB显存,性能反升4%
Compression Strategies for Efficient Multimodal LLMs in Medical Contexts
- 按层选择策略实现结构化剪枝,提升压缩效率
- 激活感知量化配合剪枝流程,显存减少70%
- 适合资源受限场景下的医疗AI部署
多模态大语言模型在医疗领域潜力巨大,但计算开销高,需高效压缩技术。本文评估了结构化剪枝与激活感知量化对微调后的LLAVA模型在医疗应用中的影响。提出一种新的剪枝层选择方法,分析不同量化技术,并在剪枝-微调-量化流水线中评估性能权衡。所提方法使7B参数的MLLM在4GB显存内运行,相比传统压缩技术在相同压缩比下内存降低70%,且模型性能提升4%。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) hold huge potential for usage in the medical domain, but their computational costs necessitate efficient compression techniques. This paper evaluates the impact of structural pruning and activation-aware quantization on a fine-tuned LLAVA model for medical applications. We propose a novel layer selection method for pruning, analyze different quantization techniques, and assess the performance trade-offs in a prune-SFT-quantize pipeline. Our proposed method enables MLLMs with 7B parameters to run within 4 GB of VRAM, reducing memory usage by 70% while achieving 4% higher model performance compared to traditional pruning and quantization techniques in the same compression ratio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。