arXiv:2510.12444cs.CV2025-10综述被引 3

首篇系统综述纵向胸片报告生成,解析数据、模型与评估方法

A Review of Longitudinal Radiology Report Generation: Dataset Composition, Methods, and Performance Evaluation

  • 梳理纵向数据构建策略与适配架构设计
  • 揭示时间序列信息对模型性能的关键作用
  • 适合医疗AI研究者和临床辅助系统开发者

胸部X光是现代医学中广泛应用的诊断工具,其高使用率给放射科医生带来巨大工作负担。为缓解这一问题,视觉语言模型被用于自动化胸部X光报告生成(CXRRRG),旨在实现临床准确描述并减少人工工作量。然而,传统方法多基于单张图像,无法捕捉纵向上下文,难以生成具有临床意义的对比性描述。近年来,越来越多研究关注将纵向数据融入报告生成,使模型能像放射科医生一样利用历史影像进行诊断。但现有综述主要聚焦单图生成,对纵向场景缺乏系统指导,导致研究者缺乏明确的设计框架。为此,本文首次提供纵向放射学报告生成(LRRG)的全面综述,涵盖数据集构建策略、报告生成架构及纵向定制化设计,以及包含纵向特异性指标与通用基准的评估协议。我们进一步总结各类方法的性能表现,并通过消融实验分析,凸显纵向信息与架构选择对模型性能的关键影响。最后,归纳当前研究的五大局限,并提出未来发展方向,旨在为该新兴领域奠定基础。

原文摘要 · Abstract (English)

Chest Xray imaging is a widely used diagnostic tool in modern medicine, and its high utilization creates substantial workloads for radiologists. To alleviate this burden, vision language models are increasingly applied to automate Chest Xray radiology report generation (CXRRRG), aiming for clinically accurate descriptions while reducing manual effort. Conventional approaches, however, typically rely on single images, failing to capture the longitudinal context necessary for producing clinically faithful comparison statements. Recently, growing attention has been directed toward incorporating longitudinal data into CXR RRG, enabling models to leverage historical studies in ways that mirror radiologists diagnostic workflows. Nevertheless, existing surveys primarily address single image CXRRRG and offer limited guidance for longitudinal settings, leaving researchers without a systematic framework for model design. To address this gap, this survey provides the first comprehensive review of longitudinal radiology report generation (LRRG). Specifically, we examine dataset construction strategies, report generation architectures alongside longitudinally tailored designs, and evaluation protocols encompassing both longitudinal specific measures and widely used benchmarks. We further summarize LRRG methods performance, alongside analyses of different ablation studies, which collectively highlight the critical role of longitudinal information and architectural design choices in improving model performance. Finally, we summarize five major limitations of current research and outline promising directions for future development, aiming to lay a foundation for advancing this emerging field.

纵向生成医学影像报告生成视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。