让大模型当医生助手,提升临床协作效率
Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
- 设计医生工作流对齐的任务与评测集
- 构建包含9.2万条问答的中文医疗数据集
- 适合医学AI研发者与临床辅助系统开发者
大语言模型在医疗领域虽能提供临床建议,但直接面向患者使用存在安全风险。为此,本文提出将大模型定位为医生的协作助手而非患者交互对象。通过两阶段需求调研,识别真实临床流程中的痛点,进而构建了涵盖22项临床任务、27个专科的大型中文医疗数据集DoctorFLAN,包含9.2万条问答。为评估模型在医生端应用的表现,提出了DoctorFLAN-test(550条单轮问答)和DotaBench(74轮多轮对话)两个评测基准。实验对比了十余种主流大模型,结果表明DoctorFLAN显著提升了开源模型在医疗场景下的表现,促进其与医生工作流对齐,并补充现有以患者为中心的医疗模型体系。本研究为医生导向型医疗大模型的发展提供了重要资源与框架。
原文摘要 · Abstract (English)
The rise of large language models (LLMs) has transformed healthcare by offering clinical guidance, yet their direct deployment to patients poses safety risks due to limited domain expertise. To mitigate this, we propose repositioning LLMs as clinical assistants that collaborate with experienced physicians rather than interacting with patients directly. We conduct a two-stage inspiration-feedback survey to identify real-world needs in clinical workflows. Guided by this, we construct DoctorFLAN, a large-scale Chinese medical dataset comprising 92,000 Q&A instances across 22 clinical tasks and 27 specialties. To evaluate model performance in doctor-facing applications, we introduce DoctorFLAN-test (550 single-turn Q&A items) and DotaBench (74 multi-turn conversations). Experimental results with over ten popular LLMs demonstrate that DoctorFLAN notably improves the performance of open-source LLMs in medical contexts, facilitating their alignment with physician workflows and complementing existing patient-oriented models. This work contributes a valuable resource and framework for advancing doctor-centered medical LLM development
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。