arXiv:2505.01831eess.IVcs.CV2025-05中稿 · publication in Neu…被引 1

提出多尺度目标感知框架,提升眼底图像增强效果

Multi-Scale Target-Aware Representation Learning for Fundus Image Enhancement

  • 采用小波分解提取多尺度结构与细节特征
  • 在多个数据集上优于当前最优方法,且模型更轻量
  • 可聚焦病灶区域,适合眼科诊断场景

高质量眼底图像对临床筛查和眼科疾病诊断至关重要。然而,受硬件限制、操作差异和患者配合度影响,眼底图像常存在分辨率低、信噪比差的问题。近年来,眼底图像增强取得进展,但多数方法仅关注结构细节或全局特征恢复,缺乏统一的多尺度信息重建框架,且很少明确以病灶等目标为增强重点。为此,本文提出多尺度目标感知表示学习框架(MTRL-FIE)。设计多尺度特征编码器(MFE)利用小波分解嵌入低频结构信息与高频细节;构建结构保持的分层解码器(SHD),融合多尺度特征并保留局部结构平滑性;引入目标感知特征聚合(TFA)模块,强化病灶区域并抑制伪影。在多个眼底图像数据集上的实验表明,该方法在性能与泛化能力方面均优于现有技术,且模型更轻量。此外,无需微调即可推广至其他眼科图像处理任务,展现出临床应用潜力。

原文摘要 · Abstract (English)

High-quality fundus images provide essential anatomical information for clinical screening and ophthalmic disease diagnosis. Yet, due to hardware limitations, operational variability, and patient compliance, fundus images often suffer from low resolution and signal-to-noise ratio. Recent years have witnessed promising progress in fundus image enhancement. However, existing works usually focus on restoring structural details or global characteristics of fundus images, lacking a unified image enhancement framework to recover comprehensive multi-scale information. Moreover, few methods pinpoint the target of image enhancement, e.g., lesions, which is crucial for medical image-based diagnosis. To address these challenges, we propose a multi-scale target-aware representation learning framework (MTRL-FIE) for efficient fundus image enhancement. Specifically, we propose a multi-scale feature encoder (MFE) that employs wavelet decomposition to embed both low-frequency structural information and high-frequency details. Next, we design a structure-preserving hierarchical decoder (SHD) to fuse multi-scale feature embeddings for real fundus image restoration. SHD integrates hierarchical fusion and group attention mechanisms to achieve adaptive feature fusion while retaining local structural smoothness. Meanwhile, a target-aware feature aggregation (TFA) module is used to enhance pathological regions and reduce artifacts. Experimental results on multiple fundus image datasets demonstrate the effectiveness and generalizability of MTRL-FIE for fundus image enhancement. Compared to state-of-the-art methods, MTRL-FIE achieves superior enhancement performance with a more lightweight architecture. Furthermore, our approach generalizes to other ophthalmic image processing tasks without supervised fine-tuning, highlighting its potential for clinical applications.

眼底图像图像增强多尺度目标感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。