arXiv:2512.04821cs.CV2025-12被引 1

用潜在空间流匹配实现医学图像分割,生成多样结果并量化不确定性。

LatentFM: A Latent Flow Matching Approach for Generative Medical Image Segmentation

  • 在潜空间构建流匹配模型,通过变分自编码器编码图像与掩码。
  • 采样多组潜变量生成多样化分割结果,像素方差反映数据分布。
  • 输出置信图帮助医生判断预测可靠性,适合临床决策支持场景。

生成模型随着流匹配(Flow Matching, FM)的兴起取得显著进展,其无需模拟即可学习精确数据密度,在无需模拟的流基框架中展现出强大生成能力。受此启发,本文提出LatentFM,一种在潜空间进行医学图像分割的流基模型。首先设计两个变分自编码器(VAEs),将医学图像及其对应掩码编码至低维潜空间;随后估计一个基于输入图像的条件速度场以引导流演化。通过采样多个潜变量表示,方法生成具有多样性的分割输出,其像素级方差可可靠捕捉底层数据分布,实现高精度且具备不确定性感知的预测。此外,模型生成置信图以量化预测确定性,为临床分析提供更丰富信息。在ISIC-2018和CVC-Clinic两个数据集上与多种先进基线方法(包括确定性和生成式模型)对比,实验表明本方法在保持潜空间高效性的同时,获得更优的分割准确率,定性与定量评估均表现卓越。

原文摘要 · Abstract (English)

Generative models have achieved remarkable progress with the emergence of flow matching (FM). It has demonstrated strong generative capabilities and attracted significant attention as a simulation-free flow-based framework capable of learning exact data densities. Motivated by these advances, we propose LatentFM, a flow-based model operating in the latent space for medical image segmentation. To model the data distribution, we first design two variational autoencoders (VAEs) to encode both medical images and their corresponding masks into a lower-dimensional latent space. We then estimate a conditional velocity field that guides the flow based on the input image. By sampling multiple latent representations, our method synthesizes diverse segmentation outputs whose pixel-wise variance reliably captures the underlying data distribution, enabling both highly accurate and uncertainty-aware predictions. Furthermore, we generate confidence maps that quantify the model certainty, providing clinicians with richer information for deeper analysis. We conduct experiments on two datasets, ISIC-2018 and CVC-Clinic, and compare our method with several prior baselines, including both deterministic and generative approach models. Through comprehensive evaluations, both qualitative and quantitative results show that our approach achieves superior segmentation accuracy while remaining highly efficient in the latent space.

医学图像生成模型不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。