arXiv:2511.01053cs.CL2025-11被引 1

基于临床指南构建可用于评估大模型的高质量数据集

Building a Silver-Standard Dataset from NICE Guidelines for Clinical LLMs

  • 从多病种指南中提取真实患者场景与临床问题
  • 用GPT辅助生成,经验证的数据集支持模型评测
  • 适合评估医疗大模型的临床推理能力与指南遵循度

大语言模型在医疗领域应用日益广泛,但缺乏标准化的评估基准来衡量其基于指南的临床推理能力。本研究从多个诊断领域的公开指南中构建了一个经过验证的数据集。该数据集借助GPT生成,包含真实患者情境和临床问题。我们对一系列主流LLMs进行了基准测试,验证了该数据集的有效性。该框架支持对大模型的临床实用性及指南遵循度进行系统性评估。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used in healthcare, yet standardised benchmarks for evaluating guideline-based clinical reasoning are missing. This study introduces a validated dataset derived from publicly available guidelines across multiple diagnoses. The dataset was created with the help of GPT and contains realistic patient scenarios, as well as clinical questions. We benchmark a range of recent popular LLMs to showcase the validity of our dataset. The framework supports systematic evaluation of LLMs' clinical utility and guideline adherence.

医疗AI大模型评测临床指南数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。