针对MRI物理与通用电气扫描仪操作知识,构建分级评测基准,揭示大模型在实操知识上的短板。
MRI-Eval: A Tiered Benchmark for Evaluating LLM Performance on MRI Physics and GE Scanner Operations Knowledge

- 设计三级多选题体系,融合教科书、手册和专家题,覆盖1365道题目
- 大模型在通用题中准确率达93%-97%,但对通用电气操作题仅88%-95%
- 去除选项后模型表现骤降,暴露其依赖选项猜测而非真实知识
现有MRI大模型评测主要依赖教材式选择题,顶级模型得分已趋饱和,难以区分性能。尚无系统性基准评估科研实践中关键的厂商特定扫描仪操作知识。本文构建了MRI-Eval,一个分层级的评测基准,用于相对比较大模型在MRI物理与通用电气(GE)扫描仪操作知识上的表现,采用主干多选题(MCQ),辅以仅给题干和带误导提示的题干测试。该基准包含来自教科书、GE扫描仪手册、编程课程材料及专家生成的9个类别共1365道可评分题目,涵盖三个难度层级。评估了五类模型(GPT-5.4、Claude Opus 4.6、Claude Sonnet 4.6、Gemini 2.5 Pro、Llama 3.3 70B)。主评测为多选题,题干独立判断由另一LLM执行,带误导提示的题干测试应对错误用户陈述的能力。结果显示,整体多选题准确率为93.2%至97.1%;所有模型在通用电气扫描仪操作类别表现最差(88.2%至94.6%)。在仅题干测试中,前沿模型准确率降至58.4%至61.1%,而Llama 3.3 70B降至37.1%;通用电气操作题干测试准确率仅为13.8%至29.8%。结论:高多选题成绩可能掩盖自由文本回忆能力不足,尤其在厂商特定操作知识方面。MRI-Eval更适合用作相对比较基准,而非绝对能力衡量,警示不应直接使用原始大模型输出指导通用电气特定协议。
原文摘要 · Abstract (English)
Background: Existing MRI LLM benchmarks rely mainly on review-book multiple-choice questions, where top proprietary models already score highly, limiting discrimination. No systematic benchmark has evaluated vendor-specific scanner operational knowledge central to research MRI practice. Purpose: We developed MRI-Eval, a tiered benchmark for relative model comparison on MRI physics and GE scanner operations knowledge using primary multiple-choice questions (MCQ), with stem-only and primed diagnostic conditions as complementary analyses. Methods: MRI-Eval includes 1365 scored items across nine categories and three difficulty tiers from textbooks, GE scanner manuals, programming course materials, and expert-generated questions. Five model families were evaluated (GPT-5.4, Claude Opus 4.6, Claude Sonnet 4.6, Gemini 2.5 Pro, Llama 3.3 70B). MCQ was primary; stem-only removed options and used an independent LLM judge; primed stem-only tested responses to incorrect user claims. Results: Overall MCQ accuracy was 93.2% to 97.1%. GE scanner operations was the lowest category for every model (88.2% to 94.6%). In stem-only, frontier-model accuracy fell to 58.4% to 61.1%, and Llama 3.3 70B fell to 37.1%; GE scanner operations stem-only accuracy was 13.8% to 29.8%. Conclusion: High MCQ performance can mask weak free-text recall, especially for vendor-specific operational knowledge. MRI-Eval is most informative as a relative comparison benchmark rather than an absolute competency measure and supports caution in using raw LLM outputs for GE-specific protocol guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。