首个评估大模型在意大利医学入学考试表现的基准测试
MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
- 构建包含1.7万道题的意大利医学入学考试多选题集,覆盖六大学科与三难度层级
- 评估多类大模型在准确率、响应一致性(88.86%)及推理能力上的表现
- 适合意大利医疗教育AI开发者和语言模型评估研究者使用
大型语言模型在教育领域潜力巨大,但非英语专业领域的评测基准仍十分稀缺。我们提出MedBench-IT,首个针对意大利医学大学入学考试的大规模评测基准。数据源自知名备考教材出版商Edizioni Simone,包含17,410道专家编写的选择题,覆盖生物学、化学、逻辑、通识、数学、物理六大学科,分为三个难度等级。我们评估了多种模型,包括GPT-4o、Claude系列等专有大模型,以及参数量小于300亿的开源轻量级模型,关注其实际部署可行性。除准确率外,还进行严格的可复现性测试(响应一致性达88.86%,学科间略有差异)、排序偏差分析(影响极小)及推理提示有效性评估。同时考察题目可读性与模型表现的相关性,发现存在统计显著但微弱的负相关关系。MedBench-IT为意大利自然语言处理社区、教育科技开发者及实践者提供关键资源,揭示当前模型能力,并建立标准化评估方法。
原文摘要 · Abstract (English)
Large language models (LLMs) show increasing potential in education, yet benchmarks for non-English languages in specialized domains remain scarce. We introduce MedBench-IT, the first comprehensive benchmark for evaluating LLMs on Italian medical university entrance examinations. Sourced from Edizioni Simone, a leading preparatory materials publisher, MedBench-IT comprises 17,410 expert-written multiple-choice questions across six subjects (Biology, Chemistry, Logic, General Culture, Mathematics, Physics) and three difficulty levels. We evaluated diverse models including proprietary LLMs (GPT-4o, Claude series) and resource-efficient open-source alternatives (<30B parameters) focusing on practical deployability. Beyond accuracy, we conducted rigorous reproducibility tests (88.86% response consistency, varying by subject), ordering bias analysis (minimal impact), and reasoning prompt evaluation. We also examined correlations between question readability and model performance, finding a statistically significant but small inverse relationship. MedBench-IT provides a crucial resource for Italian NLP community, EdTech developers, and practitioners, offering insights into current capabilities and standardized evaluation methodology for this critical domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。