用测试时扩展提升LLM在医学影像零样本诊断中的推理能力
Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
- 通过多视角描述增强视觉语言模型的推理输入
- 测试时扩展使诊断准确率显著提升,跨模态验证有效
- 适合无标注数据场景下的临床智能辅助诊断应用
作为患者诊疗的核心,临床决策直接影响治疗结果。大型语言模型(LLM)虽表现优异,但在医学影像的基于推理的诊断任务中仍鲜有探索。由于标注成本高、数据稀缺,监督微调难以实施。本文提出一种零样本医学图像诊断框架,通过测试时扩展(test-time scaling)增强LLM在临床场景中的推理能力。给定医学图像与文本提示,视觉-语言模型生成多个视觉特征描述;这些描述被送入LLM,经测试时扩展策略融合为可靠最终诊断。我们在放射学、眼科学和组织病理学等多模态数据上验证,该方法显著提升自身及基线模型的诊断准确率。实证分析表明,第一阶段无偏提示设计提升了诊断可靠性与分类性能。
原文摘要 · Abstract (English)
As a cornerstone of patient care, clinical decision-making significantly influences patient outcomes and can be enhanced by large language models (LLMs). Although LLMs have demonstrated remarkable performance, their application to visual question answering in medical imaging, particularly for reasoning-based diagnosis, remains largely unexplored. Furthermore, supervised fine-tuning for reasoning tasks is largely impractical due to limited data availability and high annotation costs. In this work, we introduce a zero-shot framework for reliable medical image diagnosis that enhances the reasoning capabilities of LLMs in clinical settings through test-time scaling. Given a medical image and a textual prompt, a vision-language model processes a medical image along with a corresponding textual prompt to generate multiple descriptions or interpretations of visual features. These interpretations are then fed to an LLM, where a test-time scaling strategy consolidates multiple candidate outputs into a reliable final diagnosis. We evaluate our approach across various medical imaging modalities -- including radiology, ophthalmology, and histopathology -- and demonstrate that the proposed test-time scaling strategy enhances diagnostic accuracy for both our and baseline methods. Additionally, we provide an empirical analysis showing that the proposed approach, which allows unbiased prompting in the first stage, improves the reliability of LLM-generated diagnoses and enhances classification accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。