arXiv:2604.01749cs.CV2026-04中稿 · CVPR被引 4

为超声影像设计专用的图文预训练模型,提升医学理解能力

Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding

  • 构建超声专属语义标签体系与36.5万对图文数据集
  • 在分类与检索任务中超越现有方法,零样本迁移表现优异
  • 适合医学影像、AI辅助诊断研究者使用

超声成像因实时性与无辐射优势广泛应用于临床诊断。然而,现有视觉语言预训练模型(如CLIP)主要针对其他模态设计,难以直接应用于具有异质解剖结构和多样诊断特征的超声数据。为此,我们构建了包含36.5万对样本、覆盖52个解剖类别的大规模超声图像-文本数据集US-365K。建立了超声诊断分类体系(UDT),包括解剖层级分类与九维诊断属性框架(系统、器官、诊断、形态、边界、回声强度、内部特征、后方声学现象、血流)。在此基础上,提出超声-CLIP模型,引入语义软标签与语义损失以增强样本区分能力,并构建基于诊断属性文本表示的异构图结构,实现病灶-属性关系的结构化推理。大量实验显示,该方法在患者级数据划分下,在分类与检索任务中达到当前最优性能,同时在零样本、线性探测与微调任务中表现出强泛化能力。

原文摘要 · Abstract (English)

Ultrasound imaging is widely used in clinical diagnostics due to its real-time capability and radiation-free nature. However, existing vision-language pre-training models, such as CLIP, are primarily designed for other modalities, and are difficult to directly apply to ultrasound data, which exhibit heterogeneous anatomical structures and diverse diagnostic attributes. To bridge this gap, we construct US-365K, a large-scale ultrasound image-text dataset containing 365k paired samples across 52 anatomical categories. We establish Ultrasonographic Diagnostic Taxonomy (UDT) containing two hierarchical knowledge frameworks. Ultrasonographic Hierarchical Anatomical Taxonomy standardizes anatomical organization, and Ultrasonographic Diagnostic Attribute Framework formalizes nine diagnostic dimensions, including body system, organ, diagnosis, shape, margins, echogenicity, internal characteristics, posterior acoustic phenomena, and vascularity. Building upon these foundations, we propose Ultrasound-CLIP, a semantic-aware contrastive learning framework that introduces semantic soft labels and semantic loss to refine sample discrimination. Moreover, we construct a heterogeneous graph modality derived from UDAF's textual representations, enabling structured reasoning over lesion-attribute relations. Extensive experiments with patient-level data splitting demonstrate that our approach achieves state-of-the-art performance on classification and retrieval benchmarks, while also delivering strong generalization to zero-shot, linear probing, and fine-tuning tasks.

超声影像图文理解对比学习医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。