arXiv:2505.07175eess.IVcs.CV2025-05被引 18

评估医学图像生成质量的常用指标大多失效,难以发现关键解剖错误。

Metrics that matter: Evaluating image quality metrics for medical image generation

  • 用脑部MRI数据测试多种无参考图像质量指标,模拟临床常见误差。
  • 多数指标对局部解剖异常不敏感,且与下游分割任务表现相关性低。
  • 建议结合下游任务评估,谨慎使用指标,避免误判模型临床可用性。

评估合成医学影像生成模型至关重要,但极具挑战性,因临床应用需极高保真度、解剖准确性及安全性。当真实图像不可得时,通常依赖无参考图像质量指标进行评估,但其在该复杂领域的可靠性尚未明确。本研究基于脑部MRI数据(包括肿瘤和血管图像),系统评估了常用无参考指标对噪声、分布偏移及局部形态改变(模拟临床相关错误)的敏感性。进一步将指标得分与下游分割任务性能对比,分析不同生成模型架构及可控扰动下的结果。发现多数常用指标与下游任务表现相关性弱,对关键局部解剖细节极度不敏感,且可能误导性地评估分布偏移(如数据记忆化)。这揭示了误判模型成熟度的风险,可能导致存在缺陷的工具投入临床,危及患者安全。结论:确保生成模型真正适用于临床,需构建多维度验证框架,融合下游任务性能与谨慎选择的无参考指标。

原文摘要 · Abstract (English)

Evaluating generative models for synthetic medical imaging is crucial yet challenging, especially given the high standards of fidelity, anatomical accuracy, and safety required for clinical applications. Standard evaluation of generated images often relies on no-reference image quality metrics when ground truth images are unavailable, but their reliability in this complex domain is not well established. This study comprehensively assesses commonly used no-reference image quality metrics using brain MRI data, including tumour and vascular images, providing a representative exemplar for the field. We systematically evaluate metric sensitivity to a range of challenges, including noise, distribution shifts, and, critically, localised morphological alterations designed to mimic clinically relevant inaccuracies. We then compare these metric scores against model performance on a relevant downstream segmentation task, analysing results across both controlled image perturbations and outputs from different generative model architectures. Our findings reveal significant limitations: many widely-used no-reference image quality metrics correlate poorly with downstream task suitability and exhibit a profound insensitivity to localised anatomical details crucial for clinical validity. Furthermore, these metrics can yield misleading scores regarding distribution shifts, e.g. data memorisation. This reveals the risk of misjudging model readiness, potentially leading to the deployment of flawed tools that could compromise patient safety. We conclude that ensuring generative models are truly fit for clinical purpose requires a multifaceted validation framework, integrating performance on relevant downstream tasks with the cautious interpretation of carefully selected no-reference image quality metrics.

医学图像生成模型质量评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。