arXiv:2505.03401cs.CVcs.AI2025-05中稿 · IEEE Transactions …被引 11

提出动态差异感知网络,提升纵向影像报告生成中变化信息的捕捉能力。

DDaTR: Dynamic Difference-aware Temporal Residual Network for Longitudinal Radiology Report Generation

  • 设计双模块视觉编码器,分层捕获空间与时间相关性
  • 在三个数据集上优于现有方法,显著提升报告生成质量
  • 适合需要追踪病灶变化的临床场景,如肿瘤随访

放射科报告生成(RRG)可自动从医学影像生成报告,提升报告效率。纵向放射科报告生成(LRRG)进一步引入对当前与既往检查的对比能力,以跟踪临床发现的时间演变。现有方法仅使用预训练视觉编码器提取当前与既往图像特征并拼接生成报告,难以有效捕捉空间与时间相关性,导致差异信息提取不足,未能充分反映病情进展,影响性能。为此,本文提出动态差异感知时间残差网络(DDaTR)。DDaTR在视觉编码器每阶段引入两个模块:动态特征对齐模块(DFAM)用于跨模态对齐既往特征,保障既往临床信息完整性;在增强后的既往特征基础上,动态差异感知模块(DDAM)通过识别跨检查间关系,主动捕捉有利差异信息。同时,采用动态残差网络单向传递纵向信息,有效建模时间依赖关系。大量实验表明,该方法在三个基准数据集上均显著优于现有方法,在RRG和LRRG任务中均表现优异。

原文摘要 · Abstract (English)

Radiology Report Generation (RRG) automates the creation of radiology reports from medical imaging, enhancing the efficiency of the reporting process. Longitudinal Radiology Report Generation (LRRG) extends RRG by incorporating the ability to compare current and prior exams, facilitating the tracking of temporal changes in clinical findings. Existing LRRG approaches only extract features from prior and current images using a visual pre-trained encoder, which are then concatenated to generate the final report. However, these methods struggle to effectively capture both spatial and temporal correlations during the feature extraction process. Consequently, the extracted features inadequately capture the information of difference across exams and thus underrepresent the expected progressions, leading to sub-optimal performance in LRRG. To address this, we develop a novel dynamic difference-aware temporal residual network (DDaTR). In DDaTR, we introduce two modules at each stage of the visual encoder to capture multi-level spatial correlations. The Dynamic Feature Alignment Module (DFAM) is designed to align prior features across modalities for the integrity of prior clinical information. Prompted by the enriched prior features, the dynamic difference-aware module (DDAM) captures favorable difference information by identifying relationships across exams. Furthermore, our DDaTR employs the dynamic residual network to unidirectionally transmit longitudinal information, effectively modelling temporal correlations. Extensive experiments demonstrated superior performance over existing methods on three benchmarks, proving its efficacy in both RRG and LRRG tasks.

影像报告生成纵向分析差异感知残差网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。