用心理测量学方法评估医疗大模型,更真实反映其专业能力。
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
- 基于项目反应理论,同时建模模型能力和题目难度/区分度。
- 在11个医学领域评估71个模型,外推预测准确率达83.3%。
- 发现模型在不同领域能力差异大,且存在两种响应模式。
基于准确率的大型语言模型(LLM)评估仅反映特定基准的表现,而非真实的医疗专业能力:它将所有问题视为同等信息量,混淆了模型能力与题目特征,导致排名随基准选择而变化。为此,我们提出MedIRT,一种基于项目反应理论(IRT)的心理测量评估框架,(1)联合建模潜在能力与题目层面的难度和区分度,(2)包含基准完整性验证,确保每类题目测量单一、连贯的核心能力。我们前瞻性地在覆盖11个医学主题的美国医师执照考试(USMLE)对齐基准上评估了71种不同的LLM。作为内部验证,MedIRT对未见问题的模型回答预测准确率达到83.3%。作为外部验证,基于IRT的排名在6个独立外部医学基准中优于准确率排名——包括专家偏好、整体临床任务、安全判断及开放问答,实现4胜0负,方差降低18%。实质性发现显示,领域级能力图谱揭示了显著的领域特异性异质性,而聚合准确率掩盖了这一现象。作为诊断工具,难度层级分析揭示出两种截然不同的响应模式(对难度敏感型与对难度不敏感型),需采取根本不同的干预策略。这些结果确立了项目感知的心理测量评估是医学领域评估LLM更有效、更稳定的基石,对任何可验证基准完整性且题目难度和区分度有意义变化的高风险领域具有潜在影响。
原文摘要 · Abstract (English)
Accuracy-based evaluation of Large Language Models (LLMs) measures benchmark-specific performance rather than underlying medical competency: it treats all questions as equally informative, conflates model ability with item characteristics, and thereby produces rankings that vary with benchmark choice. To address this, we introduce MedIRT, a psychometric evaluation framework grounded in Item Response Theory (IRT) that (1) jointly models latent competency and item-level difficulty and discrimination, and (2) includes benchmark integrity validation to ensure items within each topic measure a single, coherent underlying ability. We prospectively evaluate 71 diverse LLMs on a USMLE-aligned benchmark across 11 medical topics. As internal validation, MedIRT correctly predicts held-out LLM responses on unseen questions with 83.3% accuracy. As external validation, IRT-based rankings outperform accuracy-based rankings across 6 independent external medical benchmarks -- including expert preferences, holistic clinical tasks, safety judgments, and open-ended queries -- achieving 4 wins, 0 losses, and 18% lower variance. As a substantive finding, topic-level competency profiles expose striking domain-specific heterogeneity that aggregate accuracy masks. As a diagnostic tool, difficulty-tier analysis reveals two distinct response profiles (difficulty-sensitive responding and difficulty-insensitive responding) that require fundamentally different interventions. These results establish item-aware psychometric evaluation as a more valid and stable foundation for assessing LLMs in medicine, with potential implications for any high-stakes domain where benchmark integrity can be validated, and items vary meaningfully in difficulty and discrimination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。