arXiv:2509.23054cs.CV2025-09

用文本引导的可控制掩码提升医学图像自监督学习效果

Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis

  • 根据文本提示定位关键区域,有选择地掩码病灶相关区域
  • 在多种医学影像上实现最高3.1%分类准确率提升,掩码比例低至40%
  • 适合需要精准定位与少标注数据的医疗视觉任务

医学影像领域标注数据稀缺,制约了视觉模型的训练。自监督掩码图像建模(MIM)虽具潜力,但现有方法多依赖随机高比例掩码,效率低且语义对齐差。部分区域感知方法依赖重建启发式或监督信号,泛化能力受限。本文提出Mask What Matters:一种可控文本引导的掩码框架。通过视觉语言模型进行提示驱动的区域定位,灵活区分性地掩码,突出诊断相关区域,减少背景冗余。该设计提升了语义对齐与表征学习能力,增强了跨任务泛化性。在脑部MRI、胸部CT和肺部X光等多模态数据上评估显示,本方法显著优于SparK等现有MIM方法,分类准确率最高提升+3.1个百分点,框平均精度(BoxAP)+1.3,掩码平均精度(MaskAP)+1.1;同时掩码比例大幅降低至40%,远低于基准的70%。结果表明,可控文本驱动的掩码能实现语义对齐的自监督学习,推动医学图像分析模型发展。

原文摘要 · Abstract (English)

The scarcity of annotated data in specialized domains such as medical imaging presents significant challenges to training robust vision models. While self-supervised masked image modeling (MIM) offers a promising solution, existing approaches largely rely on random high-ratio masking, leading to inefficiency and poor semantic alignment. Moreover, region-aware variants typically depend on reconstruction heuristics or supervised signals, limiting their adaptability across tasks and modalities. We propose Mask What Matters, a controllable text-guided masking framework for self-supervised medical image analysis. By leveraging vision-language models for prompt-based region localization, our method flexibly applies differentiated masking to emphasize diagnostically relevant regions while reducing redundancy in background areas. This controllable design enables better semantic alignment, improved representation learning, and stronger cross-task generalizability. Comprehensive evaluation across multiple medical imaging modalities, including brain MRI, chest CT, and lung X-ray, shows that Mask What Matters consistently outperforms existing MIM methods (e.g., SparK), achieving gains of up to +3.1 percentage points in classification accuracy, +1.3 in box average precision (BoxAP), and +1.1 in mask average precision (MaskAP) for detection. Notably, it achieves these improvements with substantially lower overall masking ratios (e.g., 40\% vs. 70\%). This work demonstrates that controllable, text-driven masking can enable semantically aligned self-supervised learning, advancing the development of robust vision models for medical image analysis.

自监督学习医学图像文本引导掩码建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。