用眼动数据提升医学影像报告生成质量,让AI更像医生思考。
Gaze2Report: Radiology Report Generation via Visual-Gaze Prompt Tuning of LLMs
- 用眼动轨迹和图神经网络生成视觉-眼动联合特征
- 报告生成质量显著提升,且推理时无需实际眼动数据
- 适合医疗AI研发者与临床辅助系统开发者
现有深度学习方法在放射科报告生成中虽提升了诊断效率,但常忽视医生的医学先验知识,导致结构化解释与疾病表现之间对齐不佳。眼动数据能反映放射科医生的视觉注意力,增强特征提取的相关性与可解释性,契合人类决策过程。然而,眼动数据在多模态融合中的复杂性及采集成本高,尤其在推理阶段缺乏眼动输入,限制了其在真实临床场景的应用。为此,我们提出 Gaze2Report 框架,通过扫描路径预测模块与图神经网络(GNN)生成联合视觉-眼动标记,结合指令与报告标记,构成多模态提示,用于微调大语言模型(LLM)的 LoRA 层以实现自回归报告生成。该框架通过眼动引导的视觉学习提升报告质量,并支持推理时实时预测扫描路径,使模型无需依赖实际眼动数据即可运行。
原文摘要 · Abstract (English)
Existing deep learning methods for radiology report generation enhance diagnostic efficiency but often overlook physician-informed medical priors. This leads to a suboptimal alignment between the structured explanations and disease manifestations. Eye gaze data provides critical insights into a radiologist's visual attention, enhancing the relevance and interpretability of extracted features while aligning with human decision-making processes. However, despite its promising potential, the integration of eye gaze information into AI-driven medical imaging workflows is impeded by challenges such as the complexity of multimodal data fusion and the high cost of gaze acquisition, particularly its absence during inference, limiting its practical applicability in real-world clinical settings. To address these issues, we introduce Gaze2Report, a framework which leverages a scanpath prediction module and Graph Neural Network (GNN) to generate joint visual-gaze tokens. Combined with instruction and report tokens, these form a multimodal prompt used to fine-tune LoRA layers of large language models (LLMs) for autoregressive report generation. Gaze2Report enhances report quality through eye-gaze-guided visual learning and incorporates on-the-fly scanpath prediction, enabling the model to operate without gaze input during inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。