arXiv:2506.19694cs.CV2025-06被引 5

用少量超声样本实现精准病灶定位与良恶性分类

UltraAD: Fine-Grained Ultrasound Anomaly Classification via Few-Shot CLIP Adaptation

  • 基于视觉语言模型,融合图像与文本嵌入提升定位精度
  • 在三个乳腺超声数据集上超越现有方法,良恶性分类更准确
  • 适合医疗影像少样本场景,尤其适用于设备差异大的超声分析

医学影像中的精确异常检测对临床决策至关重要。尽管近期基于大规模正常数据的无监督或半监督异常检测方法表现良好,但缺乏对良恶性肿瘤等细粒度差异的区分能力。此外,超声成像对设备和采集参数高度敏感,导致图像间存在显著域差距。为此,我们提出UltraAD,一种基于视觉语言模型(VLM)的方法,利用少量超声样本实现泛化异常定位与细粒度分类。为提升定位性能,先将查询视觉原型的图像级标记与可学习文本嵌入融合,生成图像引导的提示特征,并进一步整合局部标记,优化局部表示以提高准确性。针对细粒度分类,构建由少量图像样本及其对应文本描述组成的记忆库,捕捉解剖与异常特异性特征。训练时保持文本嵌入冻结,仅适配图像特征以更好对齐医学数据。UltraAD在三个乳腺超声数据集上进行了广泛评估,其病灶定位与细粒度医学分类性能均优于现有最佳方法。代码将在论文接受后发布。

原文摘要 · Abstract (English)

Precise anomaly detection in medical images is critical for clinical decision-making. While recent unsupervised or semi-supervised anomaly detection methods trained on large-scale normal data show promising results, they lack fine-grained differentiation, such as benign vs. malignant tumors. Additionally, ultrasound (US) imaging is highly sensitive to devices and acquisition parameter variations, creating significant domain gaps in the resulting US images. To address these challenges, we propose UltraAD, a vision-language model (VLM)-based approach that leverages few-shot US examples for generalized anomaly localization and fine-grained classification. To enhance localization performance, the image-level token of query visual prototypes is first fused with learnable text embeddings. This image-informed prompt feature is then further integrated with patch-level tokens, refining local representations for improved accuracy. For fine-grained classification, a memory bank is constructed from few-shot image samples and corresponding text descriptions that capture anatomical and abnormality-specific features. During training, the stored text embeddings remain frozen, while image features are adapted to better align with medical data. UltraAD has been extensively evaluated on three breast US datasets, outperforming state-of-the-art methods in both lesion localization and fine-grained medical classification. The code will be released upon acceptance.

超声分析细粒度分类少样本学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。