arXiv:2410.21348cs.CLcs.AI2024-10被引 32

梳理医学大模型评估的多模态基准数据集,助力临床应用研究。

Large Language Model Benchmarks in Medical Tasks

  • 按文本、图像、多模态分类梳理医学LLM基准数据集
  • 涵盖MIMIC-III、CheXpert等10+核心数据集,支撑诊断与报告生成
  • 适合医疗AI研究者参考,推动多模态医学智能发展

随着大语言模型在医学领域的广泛应用,使用基准数据集评估其性能变得至关重要。本文系统综述了用于医学LLM任务的多种基准数据集,覆盖文本、图像及多模态类型,聚焦电子健康记录(EHR)、医患对话、医学问答和医学图像描述等方向。数据集按模态分类,讨论其结构、意义及对诊断、报告生成、预测决策支持等临床任务的影响。关键基准包括MIMIC-III、MIMIC-IV、BioASQ、PubMedQA和CheXpert,已推动医学报告生成、临床摘要和合成数据生成等任务进展。论文总结了利用这些基准在构建多模态医学智能中面临的挑战与机遇,强调需提升语言多样性、引入结构化组学数据并发展创新合成方法。本工作为医学领域大模型应用的未来发展奠定基础,促进医学人工智能的演进。

原文摘要 · Abstract (English)

With the increasing application of large language models (LLMs) in the medical domain, evaluating these models' performance using benchmark datasets has become crucial. This paper presents a comprehensive survey of various benchmark datasets employed in medical LLM tasks. These datasets span multiple modalities including text, image, and multimodal benchmarks, focusing on different aspects of medical knowledge such as electronic health records (EHRs), doctor-patient dialogues, medical question-answering, and medical image captioning. The survey categorizes the datasets by modality, discussing their significance, data structure, and impact on the development of LLMs for clinical tasks such as diagnosis, report generation, and predictive decision support. Key benchmarks include MIMIC-III, MIMIC-IV, BioASQ, PubMedQA, and CheXpert, which have facilitated advancements in tasks like medical report generation, clinical summarization, and synthetic data generation. The paper summarizes the challenges and opportunities in leveraging these benchmarks for advancing multimodal medical intelligence, emphasizing the need for datasets with a greater degree of language diversity, structured omics data, and innovative approaches to synthesis. This work also provides a foundation for future research in the application of LLMs in medicine, contributing to the evolving field of medical artificial intelligence.

医学AI大模型评测多模态基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。