arXiv:2510.21839cs.CVcs.AI2025-10

零样本评估GPT-4o识别肺部X光片肺炎的能力,发现简洁提示效果最佳。

Evaluating ChatGPT's Performance in Classifying Pneumonia from Chest X-Ray Images

  • 使用四种提示设计,测试GPT-4o在无微调下的零样本分类能力。
  • 最简提示下准确率达74%,高于基于推理的复杂提示。
  • 适合关注AI医学影像潜力与局限的研究者参考。

本研究评估了OpenAI的gpt-4o模型在零样本设置下,无需任何微调,对胸部X光片进行正常与肺炎分类的能力。采用包含400张图像的平衡测试集(每类200张),评估了四种不同提示设计的效果,从最简指令到详细推理型提示。结果表明,简洁、聚焦特征的提示达到最高分类准确率74%,而强调推理的提示表现更差。研究显示,尽管GPT-4o在医学图像理解方面展现出初步潜力,但其诊断可靠性仍有限。未来需在视觉推理与领域适配方面持续改进,方能安全应用于临床实践。

原文摘要 · Abstract (English)

In this study, we evaluate the ability of OpenAI's gpt-4o model to classify chest X-ray images as either NORMAL or PNEUMONIA in a zero-shot setting, without any prior fine-tuning. A balanced test set of 400 images (200 from each class) was used to assess performance across four distinct prompt designs, ranging from minimal instructions to detailed, reasoning-based prompts. The results indicate that concise, feature-focused prompts achieved the highest classification accuracy of 74\%, whereas reasoning-oriented prompts resulted in lower performance. These findings highlight that while ChatGPT exhibits emerging potential for medical image interpretation, its diagnostic reliability remains limited. Continued advances in visual reasoning and domain-specific adaptation are required before such models can be safely applied in clinical practice.

医学影像大模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。