首个眼科领域综合评测基准,专为评估大模型临床推理能力设计。
BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning
- 联合13位眼科专家构建多轮审核的高质量题库
- 涵盖900道题,覆盖5个医学数据集,准确率与推理质量双评估
- 适合医疗AI研究者、模型开发者用于公平对比测试
现有眼科大模型评测范围有限且过度侧重准确率。本文提出BELO(BEnchmarking LLMs for Ophthalmology),一个通过13位眼科专家多轮审核建立的标准、全面的评估基准,用于衡量眼科临床准确性与推理质量。基于关键词匹配与微调的PubMedBERT模型,从BCSC、MedMCQA、MedQA、BioASQ和PubMedQA五个医学数据集中筛选出眼科相关多选题(MCQs)。经过多轮专家审核,剔除重复与低质题目,并由10位眼科医生完善每道题的正确答案解析,再经3位资深专家终审。为验证其有效性,对6个大模型(OpenAI o1、o3-mini、GPT-4o、DeepSeek-R1、Llama-3-8B、Gemini 1.5 Pro)进行准确率、宏平均F1及五项文本生成指标(ROUGE-L、BERTScore、BARTScore、METEOR、AlignScore)评估。另由两位眼科专家对50个随机输出进行定性评审,考察准确性、完整性和全面性。最终数据集包含900道高质量、专家审校题目,来源分布为:BCSC(260)、BioASQ(10)、MedMCQA(572)、MedQA(40)、PubMedQA(18)。公开排行榜已建立,旨在推动透明化评估。重要的是,该数据集将作为保留的评估基准,确保未来模型比较的公平性与可复现性。
原文摘要 · Abstract (English)
Current benchmarks evaluating large language models (LLMs) in ophthalmology are limited in scope and disproportionately prioritise accuracy. We introduce BELO (BEnchmarking LLMs for Ophthalmology), a standardized and comprehensive evaluation benchmark developed through multiple rounds of expert checking by 13 ophthalmologists. BELO assesses ophthalmology-related clinical accuracy and reasoning quality. Using keyword matching and a fine-tuned PubMedBERT model, we curated ophthalmology-specific multiple-choice-questions (MCQs) from diverse medical datasets (BCSC, MedMCQA, MedQA, BioASQ, and PubMedQA). The dataset underwent multiple rounds of expert checking. Duplicate and substandard questions were systematically removed. Ten ophthalmologists refined the explanations of each MCQ's correct answer. This was further adjudicated by three senior ophthalmologists. To illustrate BELO's utility, we evaluated six LLMs (OpenAI o1, o3-mini, GPT-4o, DeepSeek-R1, Llama-3-8B, and Gemini 1.5 Pro) using accuracy, macro-F1, and five text-generation metrics (ROUGE-L, BERTScore, BARTScore, METEOR, and AlignScore). In a further evaluation involving human experts, two ophthalmologists qualitatively reviewed 50 randomly selected outputs for accuracy, comprehensiveness, and completeness. BELO consists of 900 high-quality, expert-reviewed questions aggregated from five sources: BCSC (260), BioASQ (10), MedMCQA (572), MedQA (40), and PubMedQA (18). A public leaderboard has been established to promote transparent evaluation and reporting. Importantly, the BELO dataset will remain a hold-out, evaluation-only benchmark to ensure fair and reproducible comparisons of future models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。