用医学知识体系自动提炼高质量生物医学训练数据,提升大模型表现
Knowledge Hierarchy Guided Biological-Medical Dataset Distillation for Domain LLM Training
- 基于医学主题词表构建知识引导的数据蒸馏框架
- 生成数据使Llama3-70B超越参数更多的GPT-4表现
- 适合生物医学大模型训练者与数据构建研究者
大型语言模型在生物医学领域的快速发展凸显了其潜力与现有开源标注文本数据集规模有限、质量参差之间的差距。此外,生物医学知识体系的内在复杂性也严重制约了这一差距的弥合。本研究探讨:大语言模型能否自身成为克服该局限的关键?为此,我们提出一种自动化框架,从海量科学文献中蒸馏高质量文本训练数据。该方法通过医学主题词表(MeSH)引导,自评估并生成更贴近生物医学领域的问答问题,建立无需人工干预的全自动工作流。我们通过全面实验评估了该框架生成数据对不同规模下游语言模型的影响。结果表明,相比生命科学领域预训练模型及代表性的闭源模型GPT-4,我们的方法显著提升了问答任务表现。尤为关键的是,生成的AI就绪数据集使基线的Llama3-70B模型在使用MedPrompt时,性能超越参数量多出数倍的GPT-4。详细案例分析与消融实验验证了框架各组件的重要性。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) in biological-medical applications has highlighted a gap between their potential and the limited scale and often low quality of available open-source annotated textual datasets. In addition, the inherent complexity of the biomedical knowledge hierarchy significantly hampers efforts to bridge this gap.Can LLMs themselves play a pivotal role in overcoming this limitation? Motivated by this question, we investigate this challenge in the present study.We propose a framework that automates the distillation of high-quality textual training data from the extensive scientific literature. Our approach self-evaluates and generates questions that are more closely aligned with the biomedical domain, guided by the biomedical knowledge hierarchy through medical subject headings (MeSH). This comprehensive framework establishes an automated workflow, thereby eliminating the need for manual intervention. Furthermore, we conducted comprehensive experiments to evaluate the impact of our framework-generated data on downstream language models of varying sizes. Our approach substantially improves question-answering tasks compared to pre-trained models from the life sciences domain and powerful close-source models represented by GPT-4. Notably, the generated AI-Ready dataset enabled the Llama3-70B base model to outperform GPT-4 using MedPrompt with multiple times the number of parameters. Detailed case studies and ablation experiments underscore the significance of each component within our framework
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。