arXiv:2511.04016cs.CV2025-11

专为胸部影像设计的大型视觉模型,显著提升诊断任务表现。

MedDChest: A Content-Aware Multimodal Foundational Vision Model for Thoracic Imaging

  • 从头预训练,使用超120万张跨模态胸部影像数据
  • 新提出内容感知裁剪策略,聚焦解剖关键区域
  • 在多种诊断任务上超越主流ImageNet模型,适合医疗图像研究

医学影像中视觉模型的表现常受限于基于域外自然图像预训练的主干网络。为解决这一根本性领域差距,我们提出MedDChest,一种专为胸部影像优化的新型基础视觉变压器(ViT)模型。该模型在包含10个公开来源、超过120万张图像的多模态数据集上从头预训练,涵盖胸片和计算机断层扫描(CT)等多种模态。核心技术创新在于提出引导式随机缩放裁剪(Guided Random Resized Crops),一种内容感知的数据增强策略,使采样偏向解剖学相关区域,克服了标准裁剪在医学影像中的低效问题。通过在多样下游诊断任务上微调验证模型有效性,实验表明MedDChest显著优于现有公开的ImageNet预训练模型。结果证实大规模域内预训练结合领域特定增强策略的优越性,为众多胸部诊断任务提供了强大且稳健的特征提取起点。模型权重将公开,以推动未来研究与应用。

原文摘要 · Abstract (English)

The performance of vision models in medical imaging is often hindered by the prevailing paradigm of fine-tuning backbones pre-trained on out-of-domain natural images. To address this fundamental domain gap, we propose MedDChest, a new foundational Vision Transformer (ViT) model optimized specifically for thoracic imaging. We pre-trained MedDChest from scratch on a massive, curated, multimodal dataset of over 1.2 million images, encompassing different modalities including Chest X-ray and Computed Tomography (CT) compiled from 10 public sources. A core technical contribution of our work is Guided Random Resized Crops, a novel content-aware data augmentation strategy that biases sampling towards anatomically relevant regions, overcoming the inefficiency of standard cropping techniques on medical scans. We validate our model's effectiveness by fine-tuning it on a diverse set of downstream diagnostic tasks. Comprehensive experiments empirically demonstrate that MedDChest significantly outperforms strong, publicly available ImageNet-pretrained models. By establishing the superiority of large-scale, in-domain pre-training combined with domain-specific data augmentation, MedDChest provides a powerful and robust feature extractor that serves as a significantly better starting point for a wide array of thoracic diagnostic tasks. The model weights will be made publicly available to foster future research and applications.

医学影像视觉模型胸部诊断ViT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。