用眼动数据解码医生看X光片时的诊断意图
Interpreting Radiologist's Intention from Eye Movements in Chest X-ray Diagnosis
- 基于Transformer架构分析眼动时空特征,还原医生的搜索意图
- 在三类标注意图的数据集上,预测准确率超越现有方法
- 适合医学AI、人机交互与临床认知研究者参考
放射科医生通过眼动在胸片中导航和解读。训练有素的医生会依据潜在疾病知识,按心理清单逐项排查,其注视行为反映了明确的诊断意图。然而现有模型难以捕捉每处注视背后的动机。本文提出深度学习方法RadGazeIntent,通过处理眼动数据的时空维度,将细粒度注视特征转化为高层次的诊断意图表示。我们对现有医学眼动数据集进行重构,构建三个带意图标签的子集:RadSeq(系统性顺序搜索)、RadExplore(不确定性驱动探索)、RadHybrid(混合模式)。实验表明,RadGazeIntent在所有意图标注数据集上均优于基线方法,能有效预测医生在特定时刻关注的病灶。
原文摘要 · Abstract (English)
Radiologists rely on eye movements to navigate and interpret medical images. A trained radiologist possesses knowledge about the potential diseases that may be present in the images and, when searching, follows a mental checklist to locate them using their gaze. This is a key observation, yet existing models fail to capture the underlying intent behind each fixation. In this paper, we introduce a deep learning-based approach, RadGazeIntent, designed to model this behavior: having an intention to find something and actively searching for it. Our transformer-based architecture processes both the temporal and spatial dimensions of gaze data, transforming fine-grained fixation features into coarse, meaningful representations of diagnostic intent to interpret radiologists' goals. To capture the nuances of radiologists' varied intention-driven behaviors, we process existing medical eye-tracking datasets to create three intention-labeled subsets: RadSeq (Systematic Sequential Search), RadExplore (Uncertainty-driven Exploration), and RadHybrid (Hybrid Pattern). Experimental results demonstrate RadGazeIntent's ability to predict which findings radiologists are examining at specific moments, outperforming baseline methods across all intention-labeled datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。