arXiv:2503.00196cs.CVcs.AI2025-03被引 21

用语言控制生成高分辨率医学影像反事实图像,精准修改病灶和设备特征。

PRISM: High-Resolution & Precise Counterfactual Medical Image Generation using Language-guided Stable Diffusion

  • 基于Stable Diffusion与语言引导,生成高精度反事实医学图像。
  • 可精确移除或添加特定病灶/医疗设备,保持其他图像特征不变。
  • 提升下游分类器鲁棒性,适合临床部署与医学数据增强研究。

医学影像深度学习系统的发展受限于虚假相关性、数据不平衡及文本标注不足等问题。为应对这些挑战,需具备对医学影像复杂性强健的架构。自然图像领域视觉-语言基础模型的快速进展,促使我们思考如何将其迁移至医学影像任务。本文提出PRISM框架,利用基础模型在Stable Diffusion基础上生成高分辨率、语言引导的医学图像反事实样本。该方法在选择性修改虚假相关性(如医疗设备)和疾病特征方面展现出前所未有的精确度,可在保留其他图像特性的同时,实现特定属性的增删。通过广泛评估,验证了PRISM在反事实生成上的先进性,并推动了更鲁棒的下游分类器发展,适用于临床可部署解决方案。代码已开源:https://github.com/Amarkr1/PRISM。

原文摘要 · Abstract (English)

Developing reliable and generalizable deep learning systems for medical imaging faces significant obstacles due to spurious correlations, data imbalances, and limited text annotations in datasets. Addressing these challenges requires architectures that are robust to the unique complexities posed by medical imaging data. Rapid advancements in vision-language foundation models within the natural image domain prompt the question of how they can be adapted for medical imaging tasks. In this work, we present PRISM, a framework that leverages foundation models to generate high-resolution, language-guided medical image counterfactuals using Stable Diffusion. Our approach demonstrates unprecedented precision in selectively modifying spurious correlations (the medical devices) and disease features, enabling the removal and addition of specific attributes while preserving other image characteristics. Through extensive evaluation, we show how PRISM advances counterfactual generation and enables the development of more robust downstream classifiers for clinically deployable solutions. To facilitate broader adoption and research, we make our code publicly available at https://github.com/Amarkr1/PRISM.

医学图像反事实生成扩散模型语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。