首个生成1024x1024医学影像的视觉语言模型,精准还原细节。
Pixel Perfect MegaMed: A Megapixel-Scale Vision-Language Foundation Model for Generating High Resolution Medical Images
- 采用多尺度Transformer架构,兼顾全局解剖结构与局部细节。
- 在CheXpert数据集上实现临床可信的胸片生成,低数据下分类性能提升。
- 专为医学术语和成像模态设计,适合医疗图像合成与数据增强场景。
医学图像合成因临床所需的高分辨率与复杂细节而面临独特挑战。传统生成模型如生成对抗网络(GANs)或变分自编码器(VAEs)虽能生成高分辨率图像,却难以保留对诊断至关重要的细粒度特征。为此,我们提出Pixel Perfect MegaMed,首个可生成1024x1024分辨率医学图像的视觉语言基础模型。该模型采用专为超高清医学图像生成设计的多尺度Transformer架构,有效保持全局解剖上下文与局部图像细节。通过针对医学术语和成像模态优化的视觉-语言对齐技术,模型在前所未有的分辨率下实现了文本描述与视觉表征的精准映射。我们在CheXpert数据集上验证其能力,成功从文本提示生成临床可信的胸部X光片。除视觉质量外,这些高分辨率合成图像在下游任务中表现优异,尤其在低数据条件下用于数据增强时带来显著性能提升。代码已公开于项目网站:https://tehraninasab.github.io/pixelperfect-megamed。
原文摘要 · Abstract (English)
Medical image synthesis presents unique challenges due to the inherent complexity and high-resolution details required in clinical contexts. Traditional generative architectures such as Generative Adversarial Networks (GANs) or Variational Auto Encoder (VAEs) have shown great promise for high-resolution image generation but struggle with preserving fine-grained details that are key for accurate diagnosis. To address this issue, we introduce Pixel Perfect MegaMed, the first vision-language foundation model to synthesize images at resolutions of 1024x1024. Our method deploys a multi-scale transformer architecture designed specifically for ultra-high resolution medical image generation, enabling the preservation of both global anatomical context and local image-level details. By leveraging vision-language alignment techniques tailored to medical terminology and imaging modalities, Pixel Perfect MegaMed bridges the gap between textual descriptions and visual representations at unprecedented resolution levels. We apply our model to the CheXpert dataset and demonstrate its ability to generate clinically faithful chest X-rays from text prompts. Beyond visual quality, these high-resolution synthetic images prove valuable for downstream tasks such as classification, showing measurable performance gains when used for data augmentation, particularly in low-data regimes. Our code is accessible through the project website - https://tehraninasab.github.io/pixelperfect-megamed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。