arXiv:2508.10865cs.CVcs.AI2025-08被引 8

GPT-5系列模型在脑肿瘤MRI问答中表现中等,尚不适用于临床。

Performance of GPT-5 in Brain Tumor MRI Reasoning

  • 用多序列MRI图像和临床信息构建脑肿瘤VQA评测集,测试大模型推理能力。
  • GPT-5-mini准确率最高达44.19%,但所有模型均未达到临床可用水平。
  • 模型表现因肿瘤类型而异,无单一模型全面领先,适合研究辅助决策系统。

在神经肿瘤学中,准确区分脑肿瘤类型对治疗方案制定至关重要。近年来,大型语言模型(LLMs)的发展推动了视觉问答(VQA)方法,将图像解读与自然语言推理相结合。本研究评估了GPT-4o、GPT-5-nano、GPT-5-mini和GPT-5在基于3个脑肿瘤分割(BraTS)数据集(胶质母细胞瘤GLI、脑膜瘤MEN、脑转移瘤MET)构建的脑肿瘤VQA基准上的表现。每个病例包含多序列MRI三平面拼接图及标准化临床特征转化的结构化VQA问题。模型在零样本链式思维设置下评估视觉与推理任务的准确性。结果显示,GPT-5-mini达到最高宏平均准确率(44.19%),其次为GPT-5(43.71%)、GPT-4o(41.49%)、GPT-5-nano(35.85%)。不同肿瘤亚型间性能差异明显,无模型在所有队列中占优。结果表明,GPT-5系列模型可在结构化神经肿瘤VQA任务中实现中等准确率,但尚未达到临床应用标准。

原文摘要 · Abstract (English)

Accurate differentiation of brain tumor types on magnetic resonance imaging (MRI) is critical for guiding treatment planning in neuro-oncology. Recent advances in large language models (LLMs) have enabled visual question answering (VQA) approaches that integrate image interpretation with natural language reasoning. In this study, we evaluated GPT-4o, GPT-5-nano, GPT-5-mini, and GPT-5 on a curated brain tumor VQA benchmark derived from 3 Brain Tumor Segmentation (BraTS) datasets - glioblastoma (GLI), meningioma (MEN), and brain metastases (MET). Each case included multi-sequence MRI triplanar mosaics and structured clinical features transformed into standardized VQA items. Models were assessed in a zero-shot chain-of-thought setting for accuracy on both visual and reasoning tasks. Results showed that GPT-5-mini achieved the highest macro-average accuracy (44.19%), followed by GPT-5 (43.71%), GPT-4o (41.49%), and GPT-5-nano (35.85%). Performance varied by tumor subtype, with no single model dominating across all cohorts. These findings suggest that GPT-5 family models can achieve moderate accuracy in structured neuro-oncological VQA tasks, but not at a level acceptable for clinical use.

脑肿瘤视觉问答大模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。