arXiv:2409.05732cs.CL2024-09被引 1

用两阶段微调让多语言医疗大模型更高效准确

Towards Democratizing Multilingual Large Language Models For Medicine Through A Two-Stage Instruction Fine-tuning Approach

  • 分两阶段:先注入通用医学知识,再针对选择题任务精调
  • 在六种语言上使用超20万条高质量数据,表现媲美主流模型
  • 适合医疗AI研究者与需要多语言服务的机构使用

开源多语言医疗大模型有望服务全球不同地区的语言群体。将通用大模型适配至医疗领域通常需持续预训练,但该方法计算成本高且常不切实际。仅在特定任务上进行指令微调,可能因缺乏广泛领域知识而难以实现最优性能。为此,我们构建了两个多语言指令微调数据集MMed-IFT和MMed-IFT-MC,包含超过20万条六种语言的高质量医疗样本。提出两阶段训练范式:第一阶段使用MMed-IFT注入通用医学知识,第二阶段利用MMed-IFT-MC对任务相关的多选题进行微调。该方法在英文及多语言基准测试中均取得具有竞争力的表现,在计算效率与性能间取得良好平衡。未来计划将数据集与模型权重公开于\url{https://github.com/SpassMed/Med-Llama3}。

原文摘要 · Abstract (English)

Open-source, multilingual medical large language models (LLMs) have the potential to serve linguistically diverse populations across different regions. Adapting generic LLMs for healthcare often requires continual pretraining, but this approach is computationally expensive and sometimes impractical. Instruction fine-tuning on a specific task may not always guarantee optimal performance due to the lack of broader domain knowledge that the model needs to understand and reason effectively in diverse scenarios. To address these challenges, we introduce two multilingual instruction fine-tuning datasets, MMed-IFT and MMed-IFT-MC, containing over 200k high-quality medical samples in six languages. We propose a two-stage training paradigm: the first stage injects general medical knowledge using MMed-IFT, while the second stage fine-tunes task-specific multiple-choice questions with MMed-IFT-MC. Our method achieves competitive results on both English and multilingual benchmarks, striking a balance between computational efficiency and performance. We plan to make our dataset and model weights public at \url{https://github.com/SpassMed/Med-Llama3} in the future.

多语言医疗LLM指令微调医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。