arXiv:2605.17087cs.CV2026-05

发现医学生成数据中潜在表示难学的问题,提出新方法提升合成数据效果。

The Learnability Gap in Medical Latent Diffusion

论文配图:The Learnability Gap in Medical Latent Diffusion
图 1 · 摘自论文原文
  • 用噪声条件分类器与图像空间蒸馏改进潜在空间结构
  • 在5种自编码器上验证性能差距持续存在,提升64倍推理速度
  • 适合研究医学生成模型、数据增强和潜在空间分析的学者

使用潜在扩散模型进行生成式数据增强是缓解医学影像类别不平衡的有前景策略,但现有方法多关注感知保真度和领域特定自编码器微调,忽视更根本的瓶颈。我们识别并形式化了‘可学习性差距’:大规模预训练自编码器能忠实编码用于医学分类的判别特征,重建空间中近乎无损,但其潜在表示对分类器而言难以学习。在涵盖胸部X光、皮肤镜、计算机断层扫描和超声心动图的四个医学基准上,无论架构、初始化或超参数如何,该差距均持续存在,且医学领域微调无法弥合。为探查并部分缩小该差距,我们开发了含FiLM层的噪声条件潜在分类器及图像空间蒸馏方法,相较图像空间模型实现64倍吞吐提升和120倍内存节省,同时作为潜在空间质量诊断工具。分析表明,潜在空间结构而非保真度或领域特异性,是实现实与合成数据性能差距闭合的主要障碍。

原文摘要 · Abstract (English)

Generative data augmentation with latent diffusion models is a promising strategy for addressing class imbalance in medical imaging, yet current approaches focus on perceptual fidelity and domain-specific autoencoder fine-tuning while neglecting a more fundamental bottleneck. We identify and formalize the learnability gap: large-scale pretrained autoencoders faithfully encode discriminative features for medical classification, as evidenced by near-lossless performance in reconstruction space, yet their latent representations are structured in ways that are difficult for classifiers to learn from. Across five autoencoder families and four medical benchmarks spanning chest radiography, dermatoscopy, computed tomography, and echocardiography, we show that this gap persists regardless of architecture, initialization strategy, or hyperparameter tuning, and that medical-domain fine-tuning of the autoencoder does not close it. To probe and partially narrow the gap, we develop noise-conditioned latent classifiers with FiLM layers and image-space distillation that offer 64x throughput and 120x memory gains over image-space models while serving as diagnostic tools for latent space quality. Our analysis provides a new framework for evaluating autoencoder latent spaces and identifies their structure, rather than their fidelity or domain specificity, as the primary obstacle to closing the performance gap between real and synthetic medical training data.

医学影像生成模型潜在空间数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。