测试28个大模型在神经外科考题上的表现,发现它们易受干扰信息影响。
Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons
- 用带干扰项的题目测试大模型,评估其抗干扰能力。
- 仅6个模型通过考试,最高得分超及格线15.7%。
- 开源模型比商用模型更易受干扰,需提升鲁棒性。
神经外科学会自评题(CNS-SANS)广泛用于神经外科住院医师备考执照考试,也常被用作评估大语言模型(LLMs)神经外科知识的基准。本研究评估了28个前沿大模型在2,904道源自CNS-SANS的神经外科考题上的表现,并引入干扰框架,测试模型对含多义词干扰项的敏感性。这些干扰项包含具有临床含义但用于非临床语境的词语,以评估其对模型性能的影响。结果显示,仅有6个模型达到通过标准,顶尖模型得分高于及格线15.7%。当加入干扰项后,各架构模型准确率显著下降,降幅高达20.4%,甚至有原本通过的模型失败。通用型与医学开源模型受干扰影响更大,优于专有模型。尽管当前大模型在神经外科类考题上表现优异,但其对文本中冗余干扰信息极为脆弱。该结果凸显了开发新型抗干扰策略的必要性,以保障大模型在临床应用中的安全与可靠。
原文摘要 · Abstract (English)
The Congress of Neurological Surgeons Self-Assessment for Neurological Surgeons (CNS-SANS) questions are widely used by neurosurgical residents to prepare for written board examinations. Recently, these questions have also served as benchmarks for evaluating large language models' (LLMs) neurosurgical knowledge. This study aims to assess the performance of state-of-the-art LLMs on neurosurgery board-like questions and to evaluate their robustness to the inclusion of distractor statements. A comprehensive evaluation was conducted using 28 large language models. These models were tested on 2,904 neurosurgery board examination questions derived from the CNS-SANS. Additionally, the study introduced a distraction framework to assess the fragility of these models. The framework incorporated simple, irrelevant distractor statements containing polysemous words with clinical meanings used in non-clinical contexts to determine the extent to which such distractions degrade model performance on standard medical benchmarks. 6 of the 28 tested LLMs achieved board-passing outcomes, with the top-performing models scoring over 15.7% above the passing threshold. When exposed to distractions, accuracy across various model architectures was significantly reduced-by as much as 20.4%-with one model failing that had previously passed. Both general-purpose and medical open-source models experienced greater performance declines compared to proprietary variants when subjected to the added distractors. While current LLMs demonstrate an impressive ability to answer neurosurgery board-like exam questions, their performance is markedly vulnerable to extraneous, distracting information. These findings underscore the critical need for developing novel mitigation strategies aimed at bolstering LLM resilience against in-text distractions, particularly for safe and effective clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。