首个专为护理设计的语言模型,提升护理问答准确率
NurseLLM: The First Specialized Language Model for Nursing
- 构建多阶段数据生成流程,打造首个大规模护理选择题数据集
- 在多个护理基准上超越同类通用及医学专用模型表现
- 探索推理与多智能体协作,为护理AI应用提供新方向
大语言模型的进展已显著改变医疗系统,但在护理等专业领域仍研究不足。本文提出NurseLLM,首个面向护理领域的专用语言模型,专用于多项选择题问答任务。我们开发了多阶段数据生成管道,构建首个大规模护理选择题数据集,用于训练覆盖广泛护理主题的LLM。同时引入多个护理评估基准以实现严格评测。大量实验表明,NurseLLM在不同基准上优于同规模的通用及医学专用模型,凸显护理专用模型的重要性。最后,我们探讨了推理与多智能体协作系统在护理中的作用,揭示其未来研究与应用潜力。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have significantly transformed medical systems. However, their potential within specialized domains such as nursing remains largely underexplored. In this work, we introduce NurseLLM, the first nursing-specialized LLM tailored for multiple choice question-answering (MCQ) tasks. We develop a multi-stage data generation pipeline to build the first large scale nursing MCQ dataset to train LLMs on a broad spectrum of nursing topics. We further introduce multiple nursing benchmarks to enable rigorous evaluation. Our extensive experiments demonstrate that NurseLLM outperforms SoTA general-purpose and medical-specialized LLMs of comparable size on different benchmarks, underscoring the importance of a specialized LLM for the nursing domain. Finally, we explore the role of reasoning and multi-agent collaboration systems in nursing, highlighting their promise for future research and applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。