arXiv:2506.01841eess.IV2025-06

用大模型当临床裁判,自动评估医学图像分割是否可用。

Beyond Pixel Agreement: Large Language Models as Clinical Guardrails for Reliable Medical Image Segmentation

  • 让大模型分步推理:回忆知识、分析图像、推断解剖、综合判断。
  • 零样本下准确率达78.12%,优于专门训练的视觉模型(72.92%)。
  • 结果可解释,适合医生和研发者用于提升AI影像评估可信度。

评估AI生成的医学图像分割在临床上的可用性面临巨大挑战,传统像素匹配指标常无法捕捉真正的诊断价值。本文提出分层临床推理框架(HCR),利用大语言模型(LLMs)作为临床判据,实现零样本可靠质量评估。HCR采用多阶段结构化提示策略,引导大模型经历知识回忆、视觉特征分析、解剖推断与临床整合的详细推理过程。我们在六种不同医学影像任务的多样化数据集上验证了HCR,结果显示,使用Gemini 2.5 Flash等模型时,分类准确率达78.12%,表现媲美甚至超过专为该任务训练的视觉模型(如ResNet50,准确率72.92%)。HCR不仅提供精准的质量判断,还生成可解释的逐步推理过程。本工作展示了经适当引导的大模型可成为复杂评估的智能裁判,为医疗影像中AI的质量控制提供了更可信、更贴近临床的路径。

原文摘要 · Abstract (English)

Evaluating AI-generated medical image segmentations for clinical acceptability poses a significant challenge, as traditional pixelagreement metrics often fail to capture true diagnostic utility. This paper introduces Hierarchical Clinical Reasoner (HCR), a novel framework that leverages Large Language Models (LLMs) as clinical guardrails for reliable, zero-shot quality assessment. HCR employs a structured, multistage prompting strategy that guides LLMs through a detailed reasoning process, encompassing knowledge recall, visual feature analysis, anatomical inference, and clinical synthesis, to evaluate segmentations. We evaluated HCR on a diverse dataset across six medical imaging tasks. Our results show that HCR, utilizing models like Gemini 2.5 Flash, achieved a classification accuracy of 78.12%, performing comparably to, and in instances exceeding, dedicated vision models such as ResNet50 (72.92% accuracy) that were specifically trained for this task. The HCR framework not only provides accurate quality classifications but also generates interpretable, step-by-step reasoning for its assessments. This work demonstrates the potential of LLMs, when appropriately guided, to serve as sophisticated evaluators, offering a pathway towards more trustworthy and clinically-aligned quality control for AI in medical imaging.

医学影像大模型评估可解释性零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。