70亿参数医疗模型在低算力下表现媲美十倍大的模型。
Development and bilingual evaluation of Japanese medical large language model within reasonably low computational resources
- 基于70亿参数模型做医疗微调,适配低算力环境。
- 日英双语评测成绩达到甚至超越十倍大模型水平。
- 英语主基模型微调日文医学数据,提升跨语言性能。
大型语言模型(LLM)的成功与缩放定律推动了更大模型的普及。尤其在医疗领域,由于安全考虑,对本地运行的LLM需求日益增长。然而,目前大多数高质量开源LLM参数量达700亿,带来显著的显卡部署与运维成本。为此,我们基于最新的70亿参数模型开展医疗适配,可在低计算资源下运行。我们在日语和英语两个语言的医疗问答基准上进行对比评估,结果表明其性能达到或超过当前十倍规模的医疗专用大模型。研究发现,在日本医学数据集上微调以英语为主导的基础模型,能同时提升两种语言的表现,验证了跨语言知识迁移的有效性。我们希望本研究能缓解经济压力,为医疗机构本地化使用LLM提供可行路径。评估代码已公开于 https://github.com/stardust-coder/japanese-lm-med-harness。
原文摘要 · Abstract (English)
The recent success of large language models (LLMs) and the scaling law has led to a widespread adoption of larger models. Particularly in the healthcare industry, there is an increasing demand for locally operated LLMs due to security concerns. However, the majority of high quality open-source LLMs have a size of 70B parameters, imposing significant financial burdens on users for GPU preparation and operation. To overcome these issues, we present a medical adaptation based on the recent 7B models, which enables the operation in low computational resources. We compare the performance on medical question-answering benchmarks in two languages (Japanese and English), demonstrating that its scores reach parity with or surpass those of currently existing medical LLMs that are ten times larger. We find that fine-tuning an English-centric base model on Japanese medical dataset improves the score in both language, supporting the effect of cross-lingual knowledge transfer. We hope that this study will alleviate financial challenges, serving as a stepping stone for clinical institutions to practically utilize LLMs locally. Our evaluation code is available at https://github.com/stardust-coder/japanese-lm-med-harness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。