量化让大模型在医疗场景下本地部署成为可能
Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
- 对12个顶尖大模型进行量化评估,覆盖多种医学任务
- 量化后显存降低75%,700亿参数模型可跑在40GB消费级显卡上
- 保留领域知识与提示响应能力,适合医疗本地化应用
大型语言模型在生物医学自然语言处理中表现卓越,但其规模增长和计算需求给医疗环境带来挑战,尤其在数据隐私要求高、资源有限的场景下难以云部署。本研究系统评估了量化对12个前沿大模型(涵盖通用与生物医学专用模型)的影响,覆盖8个基准数据集上的四大任务:命名实体识别、关系抽取、多标签分类和问答。结果表明,量化可将GPU显存占用减少高达75%,同时保持模型在多样化任务中的性能表现,使70B参数模型可在40GB消费级显卡上运行。此外,领域专有知识和对高级提示方法的响应能力基本得以保留。这些发现为安全、本地部署大容量语言模型提供了重要实践指导,推动AI技术向临床实际应用转化。
原文摘要 · Abstract (English)
Large language models have demonstrated remarkable capabilities in biomedical natural language processing, yet their rapid growth in size and computational requirements present a major barrier to adoption in healthcare settings where data privacy precludes cloud deployment and resources are limited. In this study, we systematically evaluated the impact of quantization on 12 state-of-the-art large language models, including both general-purpose and biomedical-specific models, across eight benchmark datasets covering four key tasks: named entity recognition, relation extraction, multi-label classification, and question answering. We show that quantization substantially reduces GPU memory requirements-by up to 75%-while preserving model performance across diverse tasks, enabling the deployment of 70B-parameter models on 40GB consumer-grade GPUs. In addition, domain-specific knowledge and responsiveness to advanced prompting methods are largely maintained. These findings provide significant practical and guiding value, highlighting quantization as a practical and effective strategy for enabling the secure, local deployment of large yet high-capacity language models in biomedical contexts, bridging the gap between technical advances in AI and real-world clinical translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。