arXiv:2601.15891cs.CV2026-01

无需文本数据,用自监督学习训练肺部X光影像编码器

RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture

  • 基于联合嵌入预测架构,从无标注胸片中学习图像表征
  • 在报告生成任务上超越现有图文模型,性能接近最佳视觉-语言模型
  • 适合医疗影像领域研究者,尤其关注少依赖文本数据的自监督方法

视觉-语言预训练推动了医学图像表征学习的进展,但受限于成对图文数据的稀缺性和临床报告中的偏见。我们探讨是否可在无语言监督的情况下训练出具有竞争力的放射科编码器。提出RadJEPA,一种基于联合嵌入预测架构的自监督框架,在约84万张未标注胸片上预训练。该模型通过可见上下文区域预测被遮蔽目标区域的潜在表示,其目标不同于图像-文本对比预训练和DINO类自蒸馏,明确建模了表示空间中的条件结构。我们在两个数据集上评估了以冻结的Vicuna-7B解码器进行放射科报告生成的表现,并将此编码器替换至四种广泛使用的视觉-语言骨干模型(MedLLaVA、Qwen-2.5、BLIP-2和Phi-4)。同时报告疾病分类与语义分割结果。在两个数据集、四个指标上,RadJEPA在使用ViT-B/14骨干网络、分辨率224×224的情况下,达到或超过最强的仅图像和视觉-语言基线。

原文摘要 · Abstract (English)

Vision-language pretraining has driven much of the recent progress in medical image representation learning, but this paradigm is constrained by the availability of paired image-text data and by the reporting bias of clinical narratives. We ask whether competitive radiology encoders can be learned without any language supervision. We introduce RadJEPA, a self-supervised framework built on a Joint Embedding Predictive Architecture and pretrained on approximately 840K unlabeled chest X-ray images. The model learns to predict latent representations of masked target regions from a visible context region, an objective that differs from both image-text contrastive pretraining and DINO-style self-distillation by explicitly modelling conditional structure in representation space. We evaluate RadJEPA primarily on radiology report generation with a frozen Vicuna-7B decoder, and additionally substitute its encoder into four widely used vision-language backbones (MedLLaVA, Qwen-2.5, BLIP-2, and Phi-4). For completeness we also report disease classification and semantic segmentation results. Across two datasets and four metrics, RadJEPA matches or exceeds the strongest image-only and vision-language baselines while using a ViT-B/14 backbone at 224 x 224 resolution.

自监督学习医学影像视觉-语言胸片分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。