arXiv:2505.01711cs.CV2025-05被引 3

用结构化文本+医学知识增强大模型,提升胸部X光片自动解读能力

Knowledge-Augmented Language Models Interpreting Structured Chest X-Ray Findings

  • 将胸部X光图像转为结构化文本,让大语言模型直接处理
  • 在病理检测、报告生成等任务上超越现有多模态模型
  • 适合临床辅助诊断系统研发者与医学AI研究者参考

胸部X光自动解读对提升临床效率和患者护理具有重要意义。尽管多模态基础模型已有进展,但如何有效利用大语言模型(LLM)处理视觉任务仍待探索。本文提出CXR-TextInter框架,通过上游图像分析生成丰富的结构化文本表示,仅基于该文本输入驱动文本导向的LLM进行解读,并集成医学知识模块以增强临床推理。为支持训练与评估,构建了MediInstruct-CXR数据集(含结构化图像表示与多样化指令-响应对)及CXR-ClinEval基准。在该基准上的实验表明,CXR-TextInter在病理检测、报告生成和视觉问答任务中均达到当前最优性能,优于现有多模态基础模型。消融实验证明知识模块至关重要。盲评结果显示,放射科专家更偏好其输出的临床质量。本工作验证了一种医疗图像AI的新范式:当视觉信息被有效结构化并融合领域知识时,可充分释放大语言模型潜力。

原文摘要 · Abstract (English)

Automated interpretation of chest X-rays (CXR) is a critical task with the potential to significantly improve clinical workflow and patient care. While recent advances in multimodal foundation models have shown promise, effectively leveraging the full power of large language models (LLMs) for this visual task remains an underexplored area. This paper introduces CXR-TextInter, a novel framework that repurposes powerful text-centric LLMs for CXR interpretation by operating solely on a rich, structured textual representation of the image content, generated by an upstream image analysis pipeline. We augment this LLM-centric approach with an integrated medical knowledge module to enhance clinical reasoning. To facilitate training and evaluation, we developed the MediInstruct-CXR dataset, containing structured image representations paired with diverse, clinically relevant instruction-response examples, and the CXR-ClinEval benchmark for comprehensive assessment across various interpretation tasks. Extensive experiments on CXR-ClinEval demonstrate that CXR-TextInter achieves state-of-the-art quantitative performance across pathology detection, report generation, and visual question answering, surpassing existing multimodal foundation models. Ablation studies confirm the critical contribution of the knowledge integration module. Furthermore, blinded human evaluation by board-certified radiologists shows a significant preference for the clinical quality of outputs generated by CXR-TextInter. Our work validates an alternative paradigm for medical image AI, showcasing the potential of harnessing advanced LLM capabilities when visual information is effectively structured and domain knowledge is integrated.

医学影像大模型知识增强胸部X光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。