arXiv:2508.19626cs.CV2025-08

通过病变特征控制生成高保真皮肤图像,提升临床可用性。

Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model

  • 基于病变测量值与类型标签引导生成,实现位置和类型可控
  • 在7种病变类型上平均FID达0.74,优于此前最优方法6.3%
  • 适合需要定制化皮肤图像的医学影像研究与数据增强场景

真实临床皮肤图像数据有限,制约深度学习模型训练。现有合成方法常生成质量低且难以控制病变位置与类型。为此,我们提出LF-VAR模型,利用量化病变测量值与病变类型标签,指导具有临床相关性的可控皮肤图像合成。该模型采用多尺度病变聚焦的向量量化变分自编码器(VQVAE)将图像编码为离散潜在表示,实现结构化标记;再通过训练于标记序列上的视觉自回归(VAR)Transformer完成图像生成。将病变区域测量值及类型作为条件嵌入,显著提升生成图像保真度。在7种病变类型上,平均FID得分为0.74,较先前最先进方法提升6.3%。结果表明,该框架可生成高保真、临床相关的合成皮肤图像。代码已公开于https://github.com/echosun1996/LF-VAR。

原文摘要 · Abstract (English)

Skin images from real-world clinical practice are often limited, resulting in a shortage of training data for deep-learning models. While many studies have explored skin image synthesis, existing methods often generate low-quality images and lack control over the lesion's location and type. To address these limitations, we present LF-VAR, a model leveraging quantified lesion measurement scores and lesion type labels to guide the clinically relevant and controllable synthesis of skin images. It enables controlled skin synthesis with specific lesion characteristics based on language prompts. We train a multiscale lesion-focused Vector Quantised Variational Auto-Encoder (VQVAE) to encode images into discrete latent representations for structured tokenization. Then, a Visual AutoRegressive (VAR) Transformer trained on tokenized representations facilitates image synthesis. Lesion measurement from the lesion region and types as conditional embeddings are integrated to enhance synthesis fidelity. Our method achieves the best overall FID score (average 0.74) among seven lesion types, improving upon the previous state-of-the-art (SOTA) by 6.3%. The study highlights our controllable skin synthesis model's effectiveness in generating high-fidelity, clinically relevant synthetic skin images. Our framework code is available at https://github.com/echosun1996/LF-VAR.

皮肤图像合成可控生成病变控制VQVAE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。