用基础模型隐空间生成可调控眼底图像,但合成与真实数据有差距。
Evaluation of Clinically Steerable Retinal Image Generation from Foundation Model Latent Spaces

- 在基础模型隐空间中实现眼底图像可控生成
- 合成图像保留临床表型信息,下游任务表现优于传统方法
- 适合医学图像生成与模型对齐研究者阅读
医学基础模型学习了具有临床意义的表型隐表示,但其支持可控图像生成的能力仍不明确。我们在表示分词器框架下评估了四种眼底基础模型,检验基础模型隐表示中编码的人口统计学和临床信息在生成合成图像时是否得以保留。结果表明,在原始基础模型内评估时,生成的表示和图像能忠实继承表型信息,多个下游预测任务表现均优于传统隐空间扩散模型。然而,当使用真实图像训练的分类器评估时,这些优势几乎消失,揭示出此前未被关注的合成到真实表示差距。研究证明基础模型隐空间是可控眼底图像合成的强大基础,同时也凸显了需进一步对齐合成表示与真实图像分布的需求。
原文摘要 · Abstract (English)
Medical foundation models learn latent representations of clinically meaningful phenotypes, yet their ability to support controllable image generation remains largely unexplored. We evaluate four retinal foundation models within the representation tokenizer framework and examine whether demographic and clinical information encoded in latent representations from foundation models is preserved during synthetic image generation. We show that generated representations and images faithfully inherit phenotype information when evaluated within their originating foundation models, consistently outperforming conventional latent diffusion on multiple downstream prediction tasks. However, these gains largely disappear when evaluated using classifiers trained on real images, revealing a previously uncharacterised synthetic-to-real representation gap. These findings demonstrate that foundation-model latent spaces provide a powerful substrate for controllable retinal synthesis while highlighting the need to better align synthetic representations with real-image distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。