用视觉语言模型发现医学影像中隐藏的图像属性关系
Leveraging Vision-Language Foundation Models to Reveal Hidden Image-Attribute Relationships in Medical Imaging
- 微调视觉语言模型挖掘图像与属性间的潜在关联
- 生成高分辨率精准编辑图像,优于结构因果模型
- 适合医学影像数据挖掘与模型可解释性研究者
视觉语言基础模型(VLMs)在通过文本引导图像生成方面表现优异,并逐渐应用于医学影像。本文首次探究:微调后的基础模型能否揭示关键且可能未知的数据特性?在胸部X光数据集上的评估表明,相比依赖结构因果模型(SCMs)的方法,该方法生成的图像具有更高分辨率和更精确的编辑效果。首次证明微调的VLMs能揭示以往因元数据粒度不足和模型容量限制而被掩盖的隐藏数据关系。实验同时揭示了这些模型在揭示数据本质属性方面的潜力,以及在准确图像编辑中的局限性、对偏差和虚假相关性的敏感性。
原文摘要 · Abstract (English)
Vision-language foundation models (VLMs) have shown impressive performance in guiding image generation through text, with emerging applications in medical imaging. In this work, we are the first to investigate the question: 'Can fine-tuned foundation models help identify critical, and possibly unknown, data properties?' By evaluating our proposed method on a chest x-ray dataset, we show that these models can generate high-resolution, precisely edited images compared to methods that rely on Structural Causal Models (SCMs) according to numerous metrics. For the first time, we demonstrate that fine-tuned VLMs can reveal hidden data relationships that were previously obscured due to available metadata granularity and model capacity limitations. Our experiments demonstrate both the potential of these models to reveal underlying dataset properties while also exposing the limitations of fine-tuned VLMs for accurate image editing and susceptibility to biases and spurious correlations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。