让医学影像报告生成模型适应低质量X光片,提升临床实用性。
Radiology Report Generation for Low-Quality X-Ray Images

- 引入自动质量评估工具,识别低质图像并建立新基准
- 提出双环训练策略,使模型在不同画质下保持稳定表现
- 适合关注真实医疗场景落地的医生与算法研究者
视觉-语言模型(VLMs)在自动化医学影像报告生成(RRG)方面取得显著进展,但现有方法隐含假设输入图像为高质量,忽略了真实临床环境中普遍存在的噪声与伪影。因此,当前模型在处理低质图像时性能严重下降。为弥合这一差距,我们提出一种专为图像质量差异设计的鲁棒报告生成框架。首先,引入自动质量评估代理(AQAA),在MIMIC-CXR数据集中识别低质样本,并建立低质量医学影像报告生成(LRRG)基准。为应对由质量退化引起的分布偏移,提出一种新颖的双环训练策略,结合双层优化与梯度一致性机制。该方法通过在不同质量条件下对齐梯度方向,使模型学习到与质量无关的诊断特征。大量实验表明,本方法能有效缓解因图像质量下降导致的模型性能退化。代码与数据将在论文录用后公开。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have significantly advanced automated Radiology Report Generation (RRG). However, existing methods implicitly assume high-quality inputs, overlooking the noise and artifacts prevalent in real-world clinical environments. Consequently, current models exhibit severe performance degradation when processing suboptimal images. To bridge this gap, we propose a robust report generation framework explicitly designed for image quality variations. We first introduce an Automated Quality Assessment Agent (AQAA) to identify low-quality samples within the MIMIC-CXR dataset and establish the Low-quality Radiology Report Generation (LRRG) benchmark. To tackle degradation-induced shifts, we propose a novel Dual-loop Training Strategy leveraging bi-level optimization and gradient consistency. This approach ensures the model learns quality-agnostic diagnostic features by aligning gradient directions across varying quality regimes. Extensive experiments demonstrate that our approach effectively mitigates model performance degradation caused by image quality deterioration. The code and data will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。