arXiv:2503.22658eess.IVcs.AI2025-03被引 1

用Tversky指数评估生成医学图像质量,更直观可靠。

Evaluation of Machine-generated Biomedical Images via A Tally-based Similarity Measure

  • 提出基于Tversky指数的图像相似性度量方法
  • 在真实与模拟数据集上验证,结果比传统方法更符合直觉
  • 适合需要高可信度评价的生物医学图像生成场景

超分辨率、图像修复、整体图像生成、无配对风格迁移及网络约束图像重建等任务均涉及机器学习生成图像,且实际真实图像在使用时不可知。这类合成图像的质量难以进行客观量化评估,但在关键生物医学应用中,评估的可靠性至关重要。本文指出,所有图像对比本质上是相对判断而非绝对差异度量;因此,采用已被广泛验证的Tversky指数可有效评估生成图像的感知相似性。该方法在多个真实和模拟图像数据集上进行了验证,结果显示:当明确特征编码选择带来的主观性和内在缺陷后,Tversky方法能产生直观合理的评估结果,而基于深层特征空间距离汇总的传统方法则表现不佳。

原文摘要 · Abstract (English)

Super-resolution, in-painting, whole-image generation, unpaired style-transfer, and network-constrained image reconstruction each include an aspect of machine-learned image synthesis where the actual ground truth is not known at time of use. It is generally difficult to quantitatively and authoritatively evaluate the quality of synthetic images; however, in mission-critical biomedical scenarios robust evaluation is paramount. In this work, all practical image-to-image comparisons really are relative qualifications, not absolute difference quantifications; and, therefore, meaningful evaluation of generated image quality can be accomplished using the Tversky Index, which is a well-established measure for assessing perceptual similarity. This evaluation procedure is developed and then demonstrated using multiple image data sets, both real and simulated. The main result is that when the subjectivity and intrinsic deficiencies of any feature-encoding choice are put upfront, Tversky's method leads to intuitive results, whereas traditional methods based on summarizing distances in deep feature spaces do not.

图像生成医学影像评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。