arXiv:2507.09097cs.CV2025-07

用眼动视频提升大模型读片能力,让通用模型超越专业医疗模型。

RadEyeVideo: Enhancing general-domain Large Vision Language Model for chest X-ray analysis with video representations of eye gaze

  • 将放射科医生眼动数据转为视频序列,保留注视时空动态。
  • 报告生成任务性能提升24.6%,诊断平均提升15.2%。
  • 适合想提升通用模型临床能力的研究者与医疗AI开发者。

大型视觉语言模型(LVLMs)在胸部X光片(CXR)分析中表现优异。为增强人机交互,已有研究引入放射科医生的眼动信息,通常通过热图或文本提示实现。然而这些方法常忽略注视的时序顺序,而该顺序可能揭示重要检查路径。本文提出新方法RadEyeVideo,将眼动固定点数据作为视频序列输入,捕捉注视的时空动态。我们在三个具备视频输入能力的通用领域开源LVLMs上评估该方法,用于胸片报告生成与疾病诊断。当以眼动视频为提示时,模型在报告生成任务中性能最高提升24.6%,在两种任务上使用缩放评估指标后平均提升15.2%。值得注意的是,该方法使通用模型LLaVA-OneVision的表现超越了专为医学设计的MAIRA-2和CheXagent等模型。本工作表明,将领域专家知识(如眼动信息)有效融入LVLMs,可显著提升通用模型在临床任务中的能力。RadEyeVideo是迈向可扩展的人本化医疗影像分析的重要一步。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have demonstrated promising performance in chest X-ray (CXR) analysis. To enhance human-computer interaction, several studies have incorporated radiologists' eye gaze, typically through heatmaps or textual prompts. However, these methods often overlook the sequential order of eye movements, which could provide valuable insights by highlighting both the areas of interest and the order in which they are examined. In this work, we propose a novel approach called RadEyeVideo that integrates radiologists' eye-fixation data as a video sequence, capturing both the temporal and spatial dynamics of their gaze. We evaluate this method in CXR report generation and disease diagnosis using three general-domain, open-source LVLMs with video input capabilities. When prompted with eye-gaze videos, model performance improves by up to 24.6% in the report generation task and on average 15.2% for both tasks using scaled evaluation metrics. Notably, RadEyeVideo enhanced an open-domain LVLM model, LLaVA-OneVision, to surpass task-specific medical LVLMs such as MAIRA-2 and CheXagent, trained on large Chest X-ray data. This work highlights that domain expert's knowledge (eye-gaze information in this case), when effectively integrated with LVLMs, can significantly enhance general-domain models' capabilities in clinical tasks. RadEyeVideo is a step toward a scalable human-centered approach of utilizing LVLMs in medical image analytics.

医学影像眼动追踪视觉语言模型人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。