无需训练即可自动识别医学影像中的身体区域,突破依赖元数据的瓶颈。
Zero-shot System for Automatic Body Region Detection for Volumetric CT and MR Images
- 利用预训练模型和规则系统实现零样本检测,不依赖标注数据。
- 在887张异构影像上,CT与MR的加权F1分别达0.947和0.914。
- 适合缺乏标注数据但需快速部署的临床自动化场景。
可靠识别解剖体部区域是众多自动化医学影像流程的前提,但现有方法仍严重依赖不可靠的DICOM元数据。当前方案多采用监督学习,限制了其在真实场景的应用。本文探究是否可借助大型预训练基础模型中的知识,在无训练情况下实现体积化CT与MR图像的体部区域检测。我们提出并系统评估三种无训练流程:(1) 基于预训练多器官分割模型的规则驱动分割系统;(2) 由放射科医生定义规则引导的多模态大语言模型(MLLM);(3) 结合视觉输入与显式解剖证据的分割感知型MLLM。所有方法在887例异构CT与MR扫描上评估,采用人工验证的解剖区域标签。基于分割的规则系统表现最强且最稳定,CT与MR的加权F1分别为0.947与0.914,展现出跨模态与异常扫描覆盖的鲁棒性。MLLM在视觉特征明显的区域表现良好,而分割感知型MLLM暴露出根本性局限。
原文摘要 · Abstract (English)
Reliable identification of anatomical body regions is a prerequisite for many automated medical imaging workflows, yet existing solutions remain heavily dependent on unreliable DICOM metadata. Current solutions mainly use supervised learning, which limits their applicability in many real-world scenarios. In this work, we investigate whether body region detection in volumetric CT and MR images can be achieved in a fully zero-shot manner by using knowledge embedded in large pre-trained foundation models. We propose and systematically evaluate three training-free pipelines: (1) a segmentation-driven rule-based system leveraging pre-trained multi-organ segmentation models, (2) a Multimodal Large Language Model (MLLM) guided by radiologist-defined rules, and (3) a segmentation-aware MLLM that combines visual input with explicit anatomical evidence. All methods are evaluated on 887 heterogeneous CT and MR scans with manually verified anatomical region labels. The segmentation-driven rule-based approach achieves the strongest and most consistent performance, with weighted F1-scores of 0.947 (CT) and 0.914 (MR), demonstrating robustness across modalities and atypical scan coverage. The MLLM performs competitively in visually distinctive regions, while the segmentation-aware MLLM reveals fundamental limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。