arXiv:2412.06287cs.CL2024-12被引 6

首个中文儿科大模型评测数据集,覆盖12类疾病共5749题。

PediaBench: A Comprehensive Chinese Pediatric Dataset for Benchmarking Large Language Models

  • 构建包含客观与主观题的中文儿科问答数据集
  • 覆盖12类儿科疾病,含4117道客观题和1632道主观题
  • 适合作为中文医疗大模型在儿科领域的评估基准

大型语言模型(LLM)在医疗领域的兴起,凸显了对标准评测数据集以评估其问答能力的迫切需求。尽管已有多个医学问答基准数据集,但它们或涵盖多科室通用知识,或仅针对特定科室而非儿科;且部分数据集仅限客观题,无法衡量模型的生成能力。因此,现有数据集难以全面评估大模型在儿科领域的问答表现。为此,我们构建了PediaBench——首个面向中文大模型评测的儿科数据集。该数据集包含4,117道客观题和1,632道主观题,覆盖12个儿科疾病类别,采用基于难度分级的综合评分标准,全面评估模型在指令遵循、知识理解、临床案例分析等方面的能力。我们在20个开源与商用大模型上进行了广泛实验,通过深入分析结果,揭示了中文环境下大模型解答儿科问题的能力现状,并指出了改进方向。代码与数据已公开于https://github.com/ACMISLab/PediaBench。

原文摘要 · Abstract (English)

The emergence of Large Language Models (LLMs) in the medical domain has stressed a compelling need for standard datasets to evaluate their question-answering (QA) performance. Although there have been several benchmark datasets for medical QA, they either cover common knowledge across different departments or are specific to another department rather than pediatrics. Moreover, some of them are limited to objective questions and do not measure the generation capacity of LLMs. Therefore, they cannot comprehensively assess the QA ability of LLMs in pediatrics. To fill this gap, we construct PediaBench, the first Chinese pediatric dataset for LLM evaluation. Specifically, it contains 4,117 objective questions and 1,632 subjective questions spanning 12 pediatric disease groups. It adopts an integrated scoring criterion based on different difficulty levels to thoroughly assess the proficiency of an LLM in instruction following, knowledge understanding, clinical case analysis, etc. Finally, we validate the effectiveness of PediaBench with extensive experiments on 20 open-source and commercial LLMs. Through an in-depth analysis of experimental results, we offer insights into the ability of LLMs to answer pediatric questions in the Chinese context, highlighting their limitations for further improvements. Our code and data are published at https://github.com/ACMISLab/PediaBench.

儿科大模型评测中文数据集医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。