提出小波扩散变换器,提升眼底图像微动脉瘤检测准确率
WDT-MD: Wavelet Diffusion Transformers for Microaneurysm Detection in Fundus Images
- 用小波分析与扩散变换器结合,增强正常视网膜结构重建
- 在IDRiD和e-ophtha数据集上达到最优像素级和图像级检测效果
- 适合眼科AI研发者,解决误检多、复现输入等问题
微动脉瘤(MAs)是糖尿病视网膜病变(DR)最早期的特异性征象,在眼底图像中表现为小于60μm的病灶,具有高度可变的光度与形态特征,导致人工筛查既费力又易出错。尽管基于扩散的异常检测在自动化MA筛查中展现出潜力,但其临床应用受限于三大根本问题:一是模型易出现“身份映射”,无意复制输入图像;二是难以区分MAs与其他异常,导致高假阳性;三是正常特征重建不佳,影响整体性能。为此,我们提出波段扩散变换器框架(WDT-MD),包含三项创新:通过噪声编码图像条件机制,在训练中扰动图像状态,避免“身份映射”;利用修复生成伪正常模式,引入像素级监督,实现对MAs与其他异常的区分;采用结合扩散变换器全局建模能力与多尺度小波分析的小波扩散变换器架构,提升正常视网膜特征重建。在IDRiD和e-ophtha MA数据集上的全面实验表明,WDT-MD在像素级与图像级检测任务上均优于现有最先进方法。该进展为早期DR筛查提供了重要技术支持。
原文摘要 · Abstract (English)
Microaneurysms (MAs), the earliest pathognomonic signs of Diabetic Retinopathy (DR), present as sub-60 $μm$ lesions in fundus images with highly variable photometric and morphological characteristics, rendering manual screening not only labor-intensive but inherently error-prone. While diffusion-based anomaly detection has emerged as a promising approach for automated MA screening, its clinical application is hindered by three fundamental limitations. First, these models often fall prey to "identity mapping", where they inadvertently replicate the input image. Second, they struggle to distinguish MAs from other anomalies, leading to high false positives. Third, their suboptimal reconstruction of normal features hampers overall performance. To address these challenges, we propose a Wavelet Diffusion Transformer framework for MA Detection (WDT-MD), which features three key innovations: a noise-encoded image conditioning mechanism to avoid "identity mapping" by perturbing image conditions during training; pseudo-normal pattern synthesis via inpainting to introduce pixel-level supervision, enabling discrimination between MAs and other anomalies; and a wavelet diffusion Transformer architecture that combines the global modeling capability of diffusion Transformers with multi-scale wavelet analysis to enhance reconstruction of normal retinal features. Comprehensive experiments on the IDRiD and e-ophtha MA datasets demonstrate that WDT-MD outperforms state-of-the-art methods in both pixel-level and image-level MA detection. This advancement holds significant promise for improving early DR screening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。