用医疗指令微调数据训练出更通用的医学语言理解模型。
BioMistral-NLU: Towards More Generalizable Medical Language Understanding through Instruction Tuning
- 设计统一提示格式,覆盖7类医疗NLU任务
- 在6个任务上超越原模型及ChatGPT、GPT-4
- 适合需要跨任务泛化的医疗AI研究者
大型语言模型如ChatGPT通过在大规模多样化指令数据上微调,具备新任务泛化能力。但这些模型在需领域知识、细粒度文本理解与结构化数据提取的医疗自然语言理解(NLU)任务中表现不佳。为此,我们:(1)提出适用于7项重要NLU任务的统一提示格式;(2)整合多种开源医疗NLU语料,构建指令微调数据集MNLU-Instruct;(3)基于该数据集对BioMistral进行微调,得到通用医疗NLU模型BioMistral-NLU。在两个主流医疗NLU基准(BLUE和BLURB)的6项任务上以零样本方式评估,结果表明,BioMistral-NLU优于原始BioMistral及商用模型ChatGPT、GPT-4。数据集无关的提示策略与跨多样化任务的指令微调提升了模型在多样医疗任务中的泛化能力。消融实验显示,在总训练样本数不变的情况下,覆盖更广任务类型能提升下游零样本泛化性能。
原文摘要 · Abstract (English)
Large language models (LLMs) such as ChatGPT are fine-tuned on large and diverse instruction-following corpora, and can generalize to new tasks. However, those instruction-tuned LLMs often perform poorly in specialized medical natural language understanding (NLU) tasks that require domain knowledge, granular text comprehension, and structured data extraction. To bridge the gap, we: (1) propose a unified prompting format for 7 important NLU tasks, (2) curate an instruction-tuning dataset, MNLU-Instruct, utilizing diverse existing open-source medical NLU corpora, and (3) develop BioMistral-NLU, a generalizable medical NLU model, through fine-tuning BioMistral on MNLU-Instruct. We evaluate BioMistral-NLU in a zero-shot setting, across 6 important NLU tasks, from two widely adopted medical NLU benchmarks: BLUE and BLURB. Our experiments show that our BioMistral-NLU outperforms the original BioMistral, as well as the proprietary LLMs - ChatGPT and GPT-4. Our dataset-agnostic prompting strategy and instruction tuning step over diverse NLU tasks enhance LLMs' generalizability across diverse medical NLU tasks. Our ablation experiments show that instruction-tuning on a wider variety of tasks, even when the total number of training instances remains constant, enhances downstream zero-shot generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。