用智能代理自动提炼高质量生物医学语料,提升大模型问答能力。
Knowledge-Driven Agentic Scientific Corpus Distillation Framework for Biomedical Large Language Models Training
- 设计多智能体协作框架,基于医学主题词表自动提取与合成数据。
- 训练后模型在生物医学问答上超越GPT-4和Med-PaLM-2,Llama3-70B表现更优。
- 适合生物医学AI研究者、大模型训练团队使用,减少人工标注依赖。
生物医学大语言模型(LLM)的语料蒸馏旨在解决开源标注科学语料数量不足、质量不高的问题,这仍是生物医学领域有效训练大模型的瓶颈。本文提出一种知识驱动的智能体式语料蒸馏框架,专为生物医学领域大模型训练设计,以应对生物医学知识层级复杂的问题。核心是多智能体协同架构,各智能体基于医学主题词表(MeSH)层级引导,自主从海量文献中提取、合成并自评估高质量文本。该框架协同生成并优化领域特定问答对,确保覆盖全面且符合生物医学本体,同时大幅减少人工干预。实验表明,基于该框架蒸馏的数据训练的语言模型在生物医学问答任务中表现显著提升,优于多个强基准模型及先进商用模型。值得注意的是,我们的AI就绪数据集使Llama3-70B在使用MedPrompt时超越GPT-4和Med-PaLM-2,尽管后者规模更大。消融研究与案例分析进一步验证了各智能体的有效性与协同作用,凸显多智能体协作在生物医学大模型训练中的潜力。
原文摘要 · Abstract (English)
Corpus distillation for biomedical large language models (LLMs) seeks to address the pressing challenge of insufficient quantity and quality in open-source annotated scientific corpora, which remains a bottleneck for effective LLM training in biomedical research. This paper proposes a knowledge-driven, agentic framework for scientific corpus distillation, tailored explicitly for LLM training in the biomedical domain, addressing the challenge posed by the complex hierarchy of biomedical knowledge. Central to our approach is a collaborative multi-agent architecture, where specialized agents, each guided by the Medical Subject Headings (MeSH) hierarchy, work in concert to autonomously extract, synthesize, and self-evaluate high-quality textual data from vast scientific literature. This agentic framework collectively generates and refines domain-specific question-answer pairs, ensuring comprehensive coverage and consistency with biomedical ontologies while minimizing manual involvement. Extensive experimental results show that language models trained on our multi-agent distilled datasets achieve notable improvements in biomedical question-answering tasks, outperforming both strong life sciences LLM baselines and advanced proprietary models. Notably, our AI-Ready dataset enables Llama3-70B to surpass GPT-4 with MedPrompt and Med-PaLM-2, despite their larger scale. Detailed ablation studies and case analyses further validate the effectiveness and synergy of each agent within the framework, highlighting the potential of multi-agent collaboration in biomedical LLM training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。