arXiv:2410.09862eess.IV2024-10中稿 · NeurIPS被引 1

用2D眼底图像指导3D扩散模型,提升OCT图像分辨率与一致性

Conditioning 3D Diffusion Models with 2D Images: Towards Standardized OCT Volumes through En Face-Informed Super-Resolution

  • 用2D眼底图像条件化3D扩散模型,实现高分辨率体积重建
  • 切片数量提升8倍后,感知相似性指标优于插值和无条件模型
  • 适合需要标准化OCT影像的临床研究与辅助诊断场景

高维度医学影像中存在严重的各向异性问题,导致解剖与病理结构量化不一致。尤其在光学相干断层扫描(OCT)中,不同数据集、研究及临床实践中切片间距差异显著。本文提出通过条件化3D扩散模型,利用临床常规获取的2D扫视激光眼底成像(SLO)数据,将OCT体积数据标准化为更低各向异性的形式。在多中心多模态MACUSTAR数据集上进行训练与评估,当切片数量提升8倍时,本方法在感知相似性指标上优于三线性插值和无SLO条件的扩散模型。定性结果显示结构更连贯、形态更一致。该方法可提升生成决策的合理性,降低幻觉风险。本工作有望推动高质量、标准化体数据成像的发展,实现更一致的定量分析。

原文摘要 · Abstract (English)

High anisotropy in volumetric medical images can lead to the inconsistent quantification of anatomical and pathological structures. Particularly in optical coherence tomography (OCT), slice spacing can substantially vary across and within datasets, studies, and clinical practices. We propose to standardize OCT volumes to less anisotropic volumes by conditioning 3D diffusion models with en face scanning laser ophthalmoscopy (SLO) imaging data, a 2D modality already commonly available in clinical practice. We trained and evaluated on data from the multicenter and multimodal MACUSTAR study. While upsampling the number of slices by a factor of 8, our method outperforms tricubic interpolation and diffusion models without en face conditioning in terms of perceptual similarity metrics. Qualitative results demonstrate improved coherence and structural similarity. Our approach allows for better informed generative decisions, potentially reducing hallucinations. We hope this work will provide the next step towards standardized high-quality volumetric imaging, enabling more consistent quantifications.

OCT扩散模型图像超分医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。