测试智能体在零样本下区分视觉相似疾病的能力
Can Agents Distinguish Visually Hard-to-Separate Diseases in a Zero-Shot Setting? A Pilot Study
- 设计基于对比仲裁的多智能体框架,提升判别能力
- 皮肤镜数据准确率提升11个百分点,误判减少
- 适合研究医疗图像零样本诊断的学者参考
多模态大语言模型的快速发展推动了基于智能体系统的兴起。现有医学影像研究多聚焦于自动化临床流程,而本文探索一个未被充分研究但临床意义重大的场景:在零样本设置下区分视觉上难以分离的疾病。我们在两个仅依赖影像的代理诊断任务上评估代表性智能体,分别是黑色素瘤与非典型痣、肺水肿与肺炎,这两组疾病在视觉特征上高度混淆,但临床管理差异显著。我们提出一种基于对比仲裁的多智能体框架。实验结果表明,在皮肤镜数据上准确率提升11个百分点,定性样本中无依据声明减少,但整体性能仍不足以用于临床部署。我们承认人类标注固有的不确定性及缺乏临床背景,这进一步限制了向真实场景的转化。在这一受控设定下,本试点研究为零样本智能体在视觉混淆场景下的表现提供了初步见解。
原文摘要 · Abstract (English)
The rapid progress of multimodal large language models (MLLMs) has led to increasing interest in agent-based systems. While most prior work in medical imaging concentrates on automating routine clinical workflows, we study an underexplored yet clinically significant setting: distinguishing visually hard-to-separate diseases in a zero-shot setting. We benchmark representative agents on two imaging-only proxy diagnostic tasks, (1) melanoma vs. atypical nevus and (2) pulmonary edema vs. pneumonia, where visual features are highly confounded despite substantial differences in clinical management. We introduce a multi-agent framework based on contrastive adjudication. Experimental results show improved diagnostic performance (an 11-percentage-point gain in accuracy on dermoscopy data) and reduced unsupported claims on qualitative samples, although overall performance remains insufficient for clinical deployment. We acknowledge the inherent uncertainty in human annotations and the absence of clinical context, which further limit the translation to real-world settings. Within this controlled setting, this pilot study provides preliminary insights into zero-shot agent performance in visually confounded scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。