测试GPT-5在乳腺钼靶图像问答中的表现,发现其虽有潜力但离临床应用仍有差距。
Is ChatGPT-5 Ready for Mammogram VQA?
- 对比GPT-4o与GPT-5在四大数据集上的乳腺影像问答能力
- 在嵌入数据集上最高达64.5%肿块分类准确率,但整体低于人类专家
- 显示通用大模型在医学影像领域仍有优化空间,适合研究者参考
乳腺钼靶视觉问答(VQA)融合图像解读与临床推理,有潜力辅助乳腺癌筛查。我们系统评估了GPT-5系列与GPT-4o在四个公开乳腺钼靶数据集(EMBED、InBreast、CMMD、CBIS-DDSM)上的BI-RADS评估、异常检测和恶性分类任务表现。GPT-5整体表现最优,但在密度(56.8%)、畸变(52.5%)、肿块(64.5%)、钙化(63.5%)和恶性(52.8%)分类上仍低于人类专家。在InBreast上,其BI-RADS准确率为36.9%,异常检测为45.9%,恶性分类为35.0%;在CMMD上,异常检测为32.3%,恶性准确率为55.0%;在CBIS-DDSM上,分别达到69.3%、66.0%和58.2%。相比人类专家,其敏感性为63.5%,特异性为52.3%。尽管表现优于前代模型,但当前性能仍不足以支持高风险临床影像应用,需针对性领域适配与优化。
原文摘要 · Abstract (English)
Mammogram visual question answering (VQA) integrates image interpretation with clinical reasoning and has potential to support breast cancer screening. We systematically evaluated the GPT-5 family and GPT-4o model on four public mammography datasets (EMBED, InBreast, CMMD, CBIS-DDSM) for BI-RADS assessment, abnormality detection, and malignancy classification tasks. GPT-5 consistently was the best performing model but lagged behind both human experts and domain-specific fine-tuned models. On EMBED, GPT-5 achieved the highest scores among GPT variants in density (56.8%), distortion (52.5%), mass (64.5%), calcification (63.5%), and malignancy (52.8%) classification. On InBreast, it attained 36.9% BI-RADS accuracy, 45.9% abnormality detection, and 35.0% malignancy classification. On CMMD, GPT-5 reached 32.3% abnormality detection and 55.0% malignancy accuracy. On CBIS-DDSM, it achieved 69.3% BI-RADS accuracy, 66.0% abnormality detection, and 58.2% malignancy accuracy. Compared with human expert estimations, GPT-5 exhibited lower sensitivity (63.5%) and specificity (52.3%). While GPT-5 exhibits promising capabilities for screening tasks, its performance remains insufficient for high-stakes clinical imaging applications without targeted domain adaptation and optimization. However, the tremendous improvements in performance from GPT-4o to GPT-5 show a promising trend in the potential for general large language models (LLMs) to assist with mammography VQA tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。