arXiv:2505.22222cs.CVcs.CL2025-05ACL被引 4

用医生看片轨迹和标注框提升肺部X光报告生成准确率

Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation

  • 引入医生眼动轨迹与框选区域作为提示信号,增强模型定位能力
  • 使LLaVA-Med的报告质量提升9.2%,临床准确率达87.3%
  • 无需微调即可生效,适合资源有限的医疗场景

多模态大语言模型在肺部X光报告自动生成方面取得进展,但仍存在幻觉和临床错误。本文提出Look & Mark(L&M)方法,将放射科医生的眼动轨迹(Look)和边界框标注(Mark)融入提示框架。相比传统微调,L&M采用上下文学习实现显著性能提升。在多个领域模型与通用模型上测试,其使CXR-LLaVA的总体指标(A.AVG)提升1.2%,使LLaVA-Med提升9.2%;通用模型LLaVA-OV结合L&M后临床平均性能(C.AVG)达87.3%,超越专为胸部X光设计的模型。专家评估显示,每份报告的临床关键错误减少0.43处,包括误报与漏报。结果表明L&M是一种高效、可扩展的AI辅助放射学方案,适用于低资源临床环境。

原文摘要 · Abstract (English)

Recent advancements in multimodal Large Language Models (LLMs) have significantly enhanced the automation of medical image analysis, particularly in generating radiology reports from chest X-rays (CXR). However, these models still suffer from hallucinations and clinically significant errors, limiting their reliability in real-world applications. In this study, we propose Look & Mark (L&M), a novel grounding fixation strategy that integrates radiologist eye fixations (Look) and bounding box annotations (Mark) into the LLM prompting framework. Unlike conventional fine-tuning, L&M leverages in-context learning to achieve substantial performance gains without retraining. When evaluated across multiple domain-specific and general-purpose models, L&M demonstrates significant gains, including a 1.2% improvement in overall metrics (A.AVG) for CXR-LLaVA compared to baseline prompting and a remarkable 9.2% boost for LLaVA-Med. General-purpose models also benefit from L&M combined with in-context learning, with LLaVA-OV achieving an 87.3% clinical average performance (C.AVG)-the highest among all models, even surpassing those explicitly trained for CXR report generation. Expert evaluations further confirm that L&M reduces clinically significant errors (by 0.43 average errors per report), such as false predictions and omissions, enhancing both accuracy and reliability. These findings highlight L&M's potential as a scalable and efficient solution for AI-assisted radiology, paving the way for improved diagnostic workflows in low-resource clinical settings.

医学影像多模态提示工程报告生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。