arXiv:2502.14260eess.IVcs.AI2025-02被引 4

提出新基准EyeBench,让眼底图像增强更贴近临床需求。

EyeBench: A Call for More Rigorous Evaluation of Retinal Image Enhancement

  • 设计多维度临床任务评估增强效果
  • 引入专家评审与配对/非配对方法公平对比
  • 揭示现有模型在临床应用中的局限

过去十年,生成模型在眼底图像增强方面取得显著进展,但其评估仍面临挑战。现有去噪指标(如PSNR、SSIM)难以延伸至真实临床研究(如血管形态一致性)。缺乏对配对与非配对增强方法的全面评估,且缺少专家参与的评价协议。为此,我们提出新型综合基准EyeBench,具备三大特性:1)多维度临床对齐下游评估,涵盖血管分割、糖尿病视网膜病变分级、去噪泛化性、病灶分割等任务;2)医学专家指导的评估设计,构建新数据集并制定人工评估协议;3)提供深度分析洞察,帮助医疗专家选择合适模型,并揭示现有方法面临的挑战。代码已开源。

原文摘要 · Abstract (English)

Over the past decade, generative models have achieved significant success in enhancement fundus images.However, the evaluation of these models still presents a considerable challenge. A comprehensive evaluation benchmark for fundus image enhancement is indispensable for three main reasons: 1) The existing denoising metrics (e.g., PSNR, SSIM) are hardly to extend to downstream real-world clinical research (e.g., Vessel morphology consistency). 2) There is a lack of comprehensive evaluation for both paired and unpaired enhancement methods, along with the need for expert protocols to accurately assess clinical value. 3) An ideal evaluation system should provide insights to inform future developments of fundus image enhancement. To this end, we propose a novel comprehensive benchmark, EyeBench, to provide insights that align enhancement models with clinical needs, offering a foundation for future work to improve the clinical relevance and applicability of generative models for fundus image enhancement. EyeBench has three appealing properties: 1) multi-dimensional clinical alignment downstream evaluation: In addition to evaluating the enhancement task, we provide several clinically significant downstream tasks for fundus images, including vessel segmentation, DR grading, denoising generalization, and lesion segmentation. 2) Medical expert-guided evaluation design: We introduce a novel dataset that promote comprehensive and fair comparisons between paired and unpaired methods and includes a manual evaluation protocol by medical experts. 3) Valuable insights: Our benchmark study provides a comprehensive and rigorous evaluation of existing methods across different downstream tasks, assisting medical experts in making informed choices. Additionally, we offer further analysis of the challenges faced by existing methods. The code is available at \url{https://github.com/Retinal-Research/EyeBench}

眼底图像生成模型临床评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。