arXiv:2604.11653cs.CV2026-04

构建眼动数据集,对比医生与AI看X光片时的判断差异。

GazeVaLM: A Multi-Observer Eye-Tracking Benchmark for Evaluating Clinical Realism in AI-Generated X-Rays

  • 收集16名专家在真实与生成X光片上的眼动数据
  • 发现医生和大模型在真假判断上注意力模式相似但准确率有差距
  • 适合研究医疗AI可解释性与人类认知对比

我们提出GazeVaLM,一个公开的眼动追踪数据集,用于研究胸部X光片真伪评估中的临床感知。数据集包含16名专家放射科医生对30张真实和30张扩散生成的合成胸部X光片的960条眼动记录,涵盖诊断评估与真伪分类(视觉图灵测试)两种条件。每组图像-观察者对提供原始眼动数据、注视点图、扫描路径、显著性密度图、结构化诊断标签及真伪判断。我们将协议扩展至6个先进多模态大模型,发布其在相同条件下预测的诊断结果、真伪标签及置信度,支持人类与AI在决策与不确定性层面的直接比较。进一步分析了眼动一致性、观察者间一致性,并对比了医生与大模型在诊断准确率与真伪识别能力上的表现。该数据集支持眼动建模、临床决策、人机比较、生成图像真实性评估与不确定性量化等研究。通过联合发布视觉注意力数据、临床标签与模型预测,旨在推动关于专家与AI如何感知、解读和评估医学影像的可复现研究。数据集可在https://huggingface.co/datasets/davidcwong/GazeVaLM获取。

原文摘要 · Abstract (English)

We introduce GazeVaLM, a public eye-tracking dataset for studying clinical perception during chest radiograph authenticity assessment. The dataset comprises 960 gaze recordings from 16 expert radiologists interpreting 30 real and 30 synthetic chest X-rays (generated by diffusion based generative AI) under two conditions: diagnostic assessment and real-fake classification (Visual Turing test). For each image-observer pair, we provide raw gaze samples, fixation maps, scanpaths, saliency density maps, structured diagnostic labels, and authenticity judgments. We extend the protocol to 6 state-of-the-art multimodal LLMs, releasing their predicted diagnoses, authenticity labels, and confidence scores under matched conditions - enabling direct human-AI comparison at both decision and uncertainty levels. We further provide analyses of gaze agreement, inter-observer consistency, and benchmarking of radiologists versus LLMs in diagnostic accuracy and authenticity detection. GazeVaLM supports research in gaze modeling, clinical decision-making, human-AI comparison, generative image realism assessment, and uncertainty quantification. By jointly releasing visual attention data, clinical labels, and model predictions, we aim to facilitate reproducible research on how experts and AI systems perceive, interpret, and evaluate medical images. The dataset is available at https://huggingface.co/datasets/davidcwong/GazeVaLM.

医学影像眼动追踪生成模型人机对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。