现有医学大模型评估数据集缺乏真实性和透明度,亟需更严谨的评测体系。
Evaluating LLMs in Medicine: A Call for Rigor, Transparency
- 审查主流医学问答数据集的科学性与临床相关性
- 发现多数数据集缺乏真实医疗场景和严格验证流程
- 呼吁建立统一评测框架,适合医疗AI研究者与政策制定者
目的:评估大语言模型在医学问答中的当前局限,重点考察其评估所用数据集的质量。方法:审查广泛使用的基准数据集(包括MedQA、MedMCQA、PubMedQA和MMLU),分析其严谨性、透明度及与临床场景的相关性;同时评估医学期刊中的挑战题作为无偏评估工具的潜力。结果:大多数现有数据集缺乏临床真实性、透明度和稳健验证流程。公开可获取的挑战题虽具一定优势,但受限于规模小、范围窄,且存在被大模型训练数据覆盖的风险。这些缺陷凸显了构建安全、全面、具代表性的数据集的必要性。结论:建立标准化评估框架至关重要。需机构与政策制定者协同合作,确保数据集与方法学具备严谨性、无偏见并反映临床复杂性。
原文摘要 · Abstract (English)
Objectives: To evaluate the current limitations of large language models (LLMs) in medical question answering, focusing on the quality of datasets used for their evaluation. Materials and Methods: Widely-used benchmark datasets, including MedQA, MedMCQA, PubMedQA, and MMLU, were reviewed for their rigor, transparency, and relevance to clinical scenarios. Alternatives, such as challenge questions in medical journals, were also analyzed to identify their potential as unbiased evaluation tools. Results: Most existing datasets lack clinical realism, transparency, and robust validation processes. Publicly available challenge questions offer some benefits but are limited by their small size, narrow scope, and exposure to LLM training. These gaps highlight the need for secure, comprehensive, and representative datasets. Conclusion: A standardized framework is critical for evaluating LLMs in medicine. Collaborative efforts among institutions and policymakers are needed to ensure datasets and methodologies are rigorous, unbiased, and reflective of clinical complexities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。