首个公开的CT眼动数据集,助力三维扫描路径建模
CT-ScanGaze: A Dataset and Baselines for 3D Volumetric Scanpath Modeling
- 设计3D扫描路径预测模型CT-Searcher,处理三维CT体积数据
- 构建首个公开CT眼动数据集CT-ScanGaze,含真实放射科医生注视序列
- 通过2D数据转3D预训练,提升模型在3D医学图像上的表现
理解放射科医生在阅读计算机断层扫描(CT)时的眼动行为,对开发可解释的计算机辅助诊断系统至关重要。然而,该领域的研究受限于缺乏公开的眼动追踪数据集以及CT体积的三维复杂性。为解决这些问题,我们提出了首个公开的CT眼动数据集CT-ScanGaze。随后,我们设计了CT-Searcher,一种专为处理CT体积而生的新型3D扫描路径预测模型,可生成类放射科医生的三维注视序列,克服了现有模型仅支持2D输入的局限。考虑到深度学习模型受益于预训练,我们开发了一套将现有2D眼动数据转换为3D眼动数据的管道,用于预训练CT-Searcher。通过对CT-ScanGaze的定性和定量评估,我们验证了方法的有效性,并为医学影像中的3D扫描路径预测提供了全面的评估框架。
原文摘要 · Abstract (English)
Understanding radiologists' eye movement during Computed Tomography (CT) reading is crucial for developing effective interpretable computer-aided diagnosis systems. However, CT research in this area has been limited by the lack of publicly available eye-tracking datasets and the three-dimensional complexity of CT volumes. To address these challenges, we present the first publicly available eye gaze dataset on CT, called CT-ScanGaze. Then, we introduce CT-Searcher, a novel 3D scanpath predictor designed specifically to process CT volumes and generate radiologist-like 3D fixation sequences, overcoming the limitations of current scanpath predictors that only handle 2D inputs. Since deep learning models benefit from a pretraining step, we develop a pipeline that converts existing 2D gaze datasets into 3D gaze data to pretrain CT-Searcher. Through both qualitative and quantitative evaluations on CT-ScanGaze, we demonstrate the effectiveness of our approach and provide a comprehensive assessment framework for 3D scanpath prediction in medical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。