用频域注意力提升OCTA单次扫描的视网膜血管3D重建质量
Freqformer: Frequency-Domain Transformer for 3-D Reconstruction and Quantification of Human Retinal Vasculature
- 基于双分支架构,融合全局空间特征与可调频域增强模块
- 在血管数量、密度、长度等指标上与拼接体数据高度一致
- 2D逐层增强比3D块处理更快且效果相当,适合临床部署
目标:从单次光学相干断层扫描血管成像(OCTA)扫描中实现人眼视网膜血管的精准3D重建与定量分析。方法:提出Freqformer,一种新型基于Transformer的模型,采用双分支结构,结合捕捉全局空间上下文的Transformer层与针对自适应频率增强设计的复数域模块。模型使用单深度平面OCTA图像进行训练,以体素合并的OCTA作为真实标签。通过二维和三维图像质量指标进行定量评估,比较了逐切片增强与三维块增强的差异。同时进行三维血管定量分析。结果:Freqformer显著优于现有卷积神经网络及Transformer方法,在图像指标上表现更优。增强后的OCTA体积在血管段数、密度、长度和血流指数等指标上与合并体积高度相关,证明其定量分析可靠性。三维方法未带来图像指标或下游三维血管量化性能提升,但推理时间增加近一个数量级,支持2D逐切片增强策略。此外,Freqformer在更大视场范围扫描上表现出优异泛化能力,生成质量超过传统体素拼接方法。结论:Freqformer能可靠地从单次扫描的OCTA生成高分辨率3D视网膜微血管结构,实现与标准体素拼接方法相当的精确血管量化。
原文摘要 · Abstract (English)
Objective: To achieve accurate 3-D reconstruction and quantitative analysis of human retinal vasculature from a single optical coherence tomography angiography (OCTA) scan. Methods: We introduce Freqformer, a novel Transformer-based model featuring a dual-branch architecture that integrates a Transformer layer for capturing global spatial context with a complex-valued frequency-domain module designed for adaptive frequency enhancement. Freqformer was trained using single depth-plane OCTA images, utilizing volumetrically merged OCTA as the ground truth. Performance was evaluated quantitatively through 2-D and 3-D image quality metrics. 2-D networks and their 3-D counterparts were compared to assess the differences between enhancing volume slice by slice and enhancing it by 3-D patches. Furthermore, 3-D quantitative vascular metrics were conducted to quantify human retinal vasculature. Results: Freqformer substantially outperformed existing convolutional neural networks and Transformer-based methods, achieving superior image metrics. Importantly, the enhanced OCTA volumes show strong correlation with the merged volumes on vascular segment count, density, length, and flow index, further underscoring its reliability for quantitative vascular analysis. 3-D counterparts did not yield additional gains in image metrics or downstream 3-D vascular quantification but incurred nearly an order-of-magnitude longer inference time, supporting our 2-D slice-wise enhancement strategy. Additionally, Freqformer showed excellent generalization capability on larger field-of-view scans, surpassing the quality of conventional volumetric merging methods. Conclusion: Freqformer reliably generates high-definition 3-D retinal microvasculature from single-scan OCTA, enabling precise vascular quantification comparable to standard volumetric merging methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。