arXiv:2411.16969cs.CV2024-11CVPR被引 19

ZoomLDM可生成多尺度图像,解决大图生成中上下文丢失问题。

ZoomLDM: Latent Diffusion Model for multi-scale image generation

  • 通过自监督嵌入实现放大感知条件控制,支持不同缩放层级生成。
  • 在4096×4096像素和4倍超分辨率下保持全局一致性,生成质量领先。
  • 适用于病理图像等数据稀缺场景,也利于多实例学习任务。

扩散模型虽已革新图像生成,但在数字病理学、卫星影像等大尺寸图像领域仍受限。由于无法直接训练超大图像(如千兆像素级),现有方法多聚焦于小固定尺寸图像块的合成,但此类方法难以捕捉全局结构与宽泛上下文,影响语义准确性。为此,本文提出ZoomLDM,一种面向多尺度图像生成的扩散模型。其核心是新型放大感知条件机制,利用自监督学习(SSL)嵌入,使模型可在不同‘缩放’层级(即大图中不同尺度的固定大小块)生成图像。ZoomLDM能生成具有上下文一致性的组织病理学图像,在所有尺度上均达到当前最优生成质量,尤其在生成整幅图像缩略图这类数据稀缺场景表现突出。多尺度特性还支持计算可行的全局一致生成,最高可达4096×4096像素及4倍超分辨率。此外,从ZoomLDM提取的多尺度特征在多实例学习任务中表现出色。

原文摘要 · Abstract (English)

Diffusion models have revolutionized image generation, yet several challenges restrict their application to large-image domains, such as digital pathology and satellite imagery. Given that it is infeasible to directly train a model on 'whole' images from domains with potential gigapixel sizes, diffusion-based generative methods have focused on synthesizing small, fixed-size patches extracted from these images. However, generating small patches has limited applicability since patch-based models fail to capture the global structures and wider context of large images, which can be crucial for synthesizing (semantically) accurate samples. To overcome this limitation, we present ZoomLDM, a diffusion model tailored for generating images across multiple scales. Central to our approach is a novel magnification-aware conditioning mechanism that utilizes self-supervised learning (SSL) embeddings and allows the diffusion model to synthesize images at different 'zoom' levels, i.e., fixed-size patches extracted from large images at varying scales. ZoomLDM synthesizes coherent histopathology images that remain contextually accurate and detailed at different zoom levels, achieving state-of-the-art image generation quality across all scales and excelling in the data-scarce setting of generating thumbnails of entire large images. The multi-scale nature of ZoomLDM unlocks additional capabilities in large image generation, enabling computationally tractable and globally coherent image synthesis up to $4096 \times 4096$ pixels and $4\times$ super-resolution. Additionally, multi-scale features extracted from ZoomLDM are highly effective in multiple instance learning experiments.

扩散模型多尺度生成图像合成病理图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。