arXiv:2602.19322cs.CVcs.AI2026-02被引 4

用稳定教师预测潜在表示,提升超声图像的自监督学习效果。

US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound

  • 采用冻结教师模型提供稳定潜在目标,避免在线更新的不稳定性。
  • 在UltraBench上分类任务表现优于或媲美主流视觉基础模型。
  • 首次系统对比多个超声领域基础模型,验证方法有效性。

超声成像因固有的噪声采集过程,给表征学习带来挑战。低信噪比和随机斑点模式使依赖像素级重建的自监督方法难以奏效。联合嵌入预测架构(JEPAs)通过预测被遮蔽的潜在表示来克服此问题。然而,标准方法依赖于对超参数敏感且计算昂贵的在线教师模型,该模型通过指数移动平均更新。本文提出US-JEPA,采用静态教师异构潜在训练(SALT)目标。通过使用冻结的、领域特定的教师模型提供稳定的潜在目标,US-JEPA实现了学生-教师优化解耦,并促使学生扩展教师的语义先验。此外,我们在UltraBench——一个涵盖多个器官和病理状况的公开数据集基准——上首次进行了所有公开可用的先进超声基础模型的严谨比较。在线性探测下,针对多样化分类任务,US-JEPA的表现与领域特定或通用视觉基础模型基线相当甚至更优。结果表明,掩码潜在表示预测为构建鲁棒的超声表征提供了稳定且高效路径。

原文摘要 · Abstract (English)

Ultrasound (US) imaging poses unique challenges for representation learning due to its inherently noisy acquisition process. The low signal-to-noise ratio and stochastic speckle patterns hinder standard self-supervised learning methods relying on a pixel-level reconstruction objective. Joint-Embedding Predictive Architectures (JEPAs) address this drawback by predicting masked latent representations rather than raw pixels. However, standard approaches depend on hyperparameter-brittle and computationally expensive online teachers updated via exponential moving average. We propose US-JEPA, a self-supervised framework that adopts the Static-teacher Asymmetric Latent Training (SALT) objective. By using a frozen, domain-specific teacher to provide stable latent targets, US-JEPA decouples student-teacher optimization and pushes the student to expand upon the semantic priors of the teacher. In addition, we provide the first rigorous comparison of all publicly available state-of-the-art ultrasound foundation models on UltraBench, a public dataset benchmark spanning multiple organs and pathological conditions. Under linear probing for diverse classification tasks, US-JEPA achieves performance competitive with or superior to domain-specific and universal vision foundation model baselines. Our results demonstrate that masked latent prediction provides a stable and efficient path toward robust ultrasound representations.

超声成像自监督学习潜在表示医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。