arXiv:2409.14988cs.CL2024-09EMNLP被引 6

连续预训练+微调可显著提升临床大模型表现,尤其混合专家模型效果更优。

Beyond Fine-tuning: Unleashing the Potential of Continuous Pretraining for Clinical LLMs

  • 用500亿词元指令微调数据增强模型临床理解能力
  • 超过2500亿词元的连续预训练使模型基础更扎实,提升后续微调效果
  • 原本为生成质量优化的NEFTune意外在临床任务中表现突出

大型语言模型(LLMs)在临床应用中展现出巨大潜力。本研究评估了四种适配临床场景的技术:连续预训练、指令微调、NEFTune和提示工程。实验基于Mistral 7B与Mixtral 8x7B模型,使用包含500亿词元的临床预训练数据集和5亿词元的指令微调数据集。评估结果表明,尽管连续预训练超过2500亿词元后收益递减,但能为指令微调打下坚实基础。值得注意的是,主要设计用于提升生成质量的NEFTune在作者基准测试中意外带来额外增益。复杂提示工程也进一步提升了性能。研究强调需根据任务特点选择微调策略,并探索创新方法以优化临床领域的大模型表现。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated significant potential in transforming clinical applications. In this study, we investigate the efficacy of four techniques in adapting LLMs for clinical use-cases: continuous pretraining, instruct fine-tuning, NEFTune, and prompt engineering. We employ these methods on Mistral 7B and Mixtral 8x7B models, leveraging a large-scale clinical pretraining dataset of 50 billion tokens and an instruct fine-tuning dataset of 500 million tokens. Our evaluation across various clinical tasks reveals the impact of each technique. While continuous pretraining beyond 250 billion tokens yields marginal improvements on its own, it establishes a strong foundation for instruct fine-tuning. Notably, NEFTune, designed primarily to enhance generation quality, surprisingly demonstrates additional gains on our benchmark. Complex prompt engineering methods further enhance performance. These findings show the importance of tailoring fine-tuning strategies and exploring innovative techniques to optimize LLM performance in the clinical domain.

临床LLM连续预训练NEFTune指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。