用医生看图时的视线数据评估医学生成图像,发现96.6%被识破造假。
Eyes Tell the Truth: GazeVal Highlights Shortcomings of Generative AI in Medical Imaging
- 结合医生眼动轨迹与诊断评估,从人类认知角度评测合成图像质量。
- 16名放射科医生识别出96.6%的顶尖生成图像为假,暴露其临床真实性缺陷。
- 适合关注医疗AI可靠性、生成图像真实性的研究者与临床开发者。
医学影像领域对高质量合成数据的需求日益增长,但现有评估多依赖计算指标,无法反映专家实际识别能力。这导致生成图像虽在数值上逼真,却缺乏临床真实性,影响AI医疗工具的可信度。为此,我们提出GazeVal框架,融合放射科医生的眼动数据与直接判读评估,从诊断和图灵测试任务中分析专家如何感知合成图像。对16位放射科医生的实验显示,采用最新生成算法的图像中,96.6%被识别为虚假,揭示了当前生成式AI在构建临床可信图像方面的显著局限。
原文摘要 · Abstract (English)
The demand for high-quality synthetic data for model training and augmentation has never been greater in medical imaging. However, current evaluations predominantly rely on computational metrics that fail to align with human expert recognition. This leads to synthetic images that may appear realistic numerically but lack clinical authenticity, posing significant challenges in ensuring the reliability and effectiveness of AI-driven medical tools. To address this gap, we introduce GazeVal, a practical framework that synergizes expert eye-tracking data with direct radiological evaluations to assess the quality of synthetic medical images. GazeVal leverages gaze patterns of radiologists as they provide a deeper understanding of how experts perceive and interact with synthetic data in different tasks (i.e., diagnostic or Turing tests). Experiments with sixteen radiologists revealed that 96.6% of the generated images (by the most recent state-of-the-art AI algorithm) were identified as fake, demonstrating the limitations of generative AI in producing clinically accurate images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。