arXiv:2410.13458cs.CL2024-10EMNLP被引 9

构建首个涵盖133个生物医学任务的指令数据集,助力大模型跨任务泛化

MedINST: Meta Dataset of Biomedical Instructions

  • 构建多领域、多任务的生物医学指令元数据集,覆盖700万+样本
  • 基于该数据集推出难度分级的MedINST32评测基准,评估模型泛化能力
  • 在多个大模型上微调后,显著提升跨任务迁移表现,适合医疗AI研究者使用

大型语言模型(LLM)在医学分析领域的应用已取得显著进展,但高质量、多样且标注完善的医学数据集仍严重匮乏。医学数据和任务在格式、规模等方面差异大,需大量预处理才能用于训练LLM。为此,我们提出MedINST,即生物医学指令元数据集,一个涵盖133个生物医学NLP任务、超过700万条训练样本的多领域、多任务指令数据集,是目前最全面的生物医学指令数据集。基于MedINST,我们构建了具有不同难度的任务的评测基准MedINST32,以评估大模型的泛化能力。我们在MedINST上微调多个大模型,并在MedINST32上进行评估,结果表明模型跨任务泛化能力得到显著增强。

原文摘要 · Abstract (English)

The integration of large language model (LLM) techniques in the field of medical analysis has brought about significant advancements, yet the scarcity of large, diverse, and well-annotated datasets remains a major challenge. Medical data and tasks, which vary in format, size, and other parameters, require extensive preprocessing and standardization for effective use in training LLMs. To address these challenges, we introduce MedINST, the Meta Dataset of Biomedical Instructions, a novel multi-domain, multi-task instructional meta-dataset. MedINST comprises 133 biomedical NLP tasks and over 7 million training samples, making it the most comprehensive biomedical instruction dataset to date. Using MedINST as the meta dataset, we curate MedINST32, a challenging benchmark with different task difficulties aiming to evaluate LLMs' generalization ability. We fine-tune several LLMs on MedINST and evaluate on MedINST32, showcasing enhanced cross-task generalization.

生物医学指令数据集大模型NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。