为眼底图像增强设计临床对齐的综合评测基准,解决评估标准不科学问题。
Bridging Restoration and Diagnosis: A Comprehensive Benchmark for Retinal Fundus Enhancement
- 引入多任务下游评估,涵盖血管分割、糖尿病视网膜病变分级等临床任务。
- 通过医学专家人工评估,精准检测病灶结构变化、背景色偏移等关键问题。
- 提供可指导模型改进的实证分析,适合临床研究者和算法开发者参考。
过去十年,生成模型在眼底图像增强方面取得进展,但其评估仍面临挑战。现有基准存在三大缺陷:(1)传统去噪指标如PSNR、SSIM无法捕捉病灶保留、血管形态一致性等临床相关特征;(2)缺乏统一协议以公平比较成对与非成对增强方法,尤其缺乏临床专家指导;(3)评估框架需提供可行动洞察,推动临床对齐模型发展。为此,我们提出EyeBench-V2,一个连接增强模型性能与临床实用性的基准。贡献包括:(1)多维度临床对齐评估:除标准指标外,新增血管分割、糖尿病视网膜病变(DR)分级、未知噪声模式泛化能力、病灶分割等任务;(2)专家引导评估设计:构建新数据集,支持成对与非成对方法公平比较,并引入医学专家结构化评估流程,重点评价病灶结构改变、背景色偏移、伪结构生成等临床敏感点;(3)可行动洞察:通过严谨的任务导向分析,为临床研究者提供决策依据,同时揭示现有方法局限,指导下一代模型设计。
原文摘要 · Abstract (English)
Over the past decade, generative models have demonstrated success in enhancing fundus images. However, the evaluation of these models remains a challenge. A benchmark for fundus image enhancement is needed for three main reasons:(1) Conventional denoising metrics such as PSNR and SSIM fail to capture clinically relevant features, such as lesion preservation and vessel morphology consistency, limiting their applicability in real-world settings; (2) There is a lack of unified evaluation protocols that address both paired and unpaired enhancement methods, particularly those guided by clinical expertise; and (3) An evaluation framework should provide actionable insights to guide future advancements in clinically aligned enhancement models. To address these gaps, we introduce EyeBench-V2, a benchmark designed to bridge the gap between enhancement model performance and clinical utility. Our work offers three key contributions:(1) Multi-dimensional clinical-alignment through downstream evaluations: Beyond standard enhancement metrics, we assess performance across clinically meaningful tasks including vessel segmentation, diabetic retinopathy (DR) grading, generalization to unseen noise patterns, and lesion segmentation. (2) Expert-guided evaluation design: We curate a novel dataset enabling fair comparisons between paired and unpaired enhancement methods, accompanied by a structured manual assessment protocol by medical experts, which evaluates clinically critical aspects such as lesion structure alterations, background color shifts, and the introduction of artificial structures. (3) Actionable insights: Our benchmark provides a rigorous, task-oriented analysis of existing generative models, equipping clinical researchers with the evidence needed to make informed decisions, while also identifying limitations in current methods to inform the design of next-generation enhancement models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。