arXiv:2603.23506cs.CLcs.AI2026-03

用自适应测试让大模型医学评估快100倍,还更准。

Leveraging Computerized Adaptive Testing for Cost-effective Evaluation of Large Language Models in Medical Benchmarking

  • 基于项目反应理论动态选题,按模型水平实时调整难度。
  • 仅用1.3%题目就达到全量测试98.8%相关性,耗时从小时级降至分钟级。
  • 适合需要快速筛查或持续监控模型医学能力的研究者使用。

大语言模型在医疗领域的广泛应用催生了高效、可靠的评估需求。传统静态基准测试成本高、易受数据污染影响,且缺乏精细的能力测量属性。本文提出并验证了一种基于项目反应理论(IRT)的计算机化自适应测试(CAT)框架,用于高效评估大模型的标准化医学知识。研究采用两阶段设计:首先通过蒙特卡洛模拟确定最优CAT配置;其次对38个大模型进行实证评估,使用人工校准的医学题库。每个模型完成全题库测试及自适应测试——后者根据实时能力估计动态选题,并在标准误≤0.3时终止。结果表明,自适应测试所得能力估计与全题库结果高度一致(相关系数r=0.988),仅使用1.3%的题目,评估时间由数小时缩短至几分钟,显著降低令牌消耗和计算成本,同时保持模型间性能排序不变。该工作建立了一个快速、低成本的大模型基础医学知识评估框架,适用于标准化预筛选与持续监测,但不替代真实临床验证或安全性前瞻性研究。

原文摘要 · Abstract (English)

The rapid proliferation of large language models (LLMs) in healthcare creates an urgent need for scalable and psychometrically sound evaluation methods. Conventional static benchmarks are costly to administer repeatedly, vulnerable to data contamination, and lack calibrated measurement properties for fine-grained performance tracking. We propose and validate a computerized adaptive testing (CAT) framework grounded in item response theory (IRT) for efficient assessment of standardized medical knowledge in LLMs. The study comprises a two-phase design: a Monte Carlo simulation to identify optimal CAT configurations and an empirical evaluation of 38 LLMs using a human-calibrated medical item bank. Each model completed both the full item bank and an adaptive test that dynamically selected items based on real-time ability estimates and terminated upon reaching a predefined reliability threshold (standard error <= 0.3). Results show that CAT-derived proficiency estimates achieved a near-perfect correlation with full-bank estimates (r = 0.988) while using only 1.3 percent of the items. Evaluation time was reduced from several hours to minutes per model, with substantial reductions in token usage and computational cost, while preserving inter-model performance rankings. This work establishes a psychometric framework for rapid, low-cost benchmarking of foundational medical knowledge in LLMs. The proposed adaptive methodology is intended as a standardized pre-screening and continuous monitoring tool and is not a substitute for real-world clinical validation or safety-oriented prospective studies.

大模型评估自适应测试医学AI效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。