用强化学习提升肺部X光报告生成质量,更准更连贯
RL-ACRGNet: Reinforcement Learning-Based Chest Radiology Report Generation Network

- 结合预训练DenseNet与多层LSTM,通过强化学习优化图文对齐
- 在IU-Xray上BLEU-4等指标提升0.47%~0.518%,优于现有模型
- 适用于临床辅助诊断,尤其适合需标准化报告的放射科场景
医学影像解读是现代临床诊断的核心,但人工撰写放射科报告耗时且易出现解释不一致。在医疗AI领域,通过深度学习自动化生成影像描述有望优化临床流程并统一诊断输出。然而,准确识别疾病和生成精确报告仍面临挑战,主要受限于细粒度视觉特征捕捉不足及临床逻辑连贯性难以保证。为此,我们提出RL-ACRGNet,一种改进的编码器-解码器模型,将预训练DenseNet编码器与多层级LSTM解码器结合,嵌入离策略强化学习框架。通过双网络结构,利用基于度量的奖励机制优化视觉-语义嵌入。实验表明,RL-ACRGNet在IU-Xray数据集上持续优于现有先进模型,BLEU-4提升0.47%、METEOR提升0.17%、ROUGE-L提升0.518%。此外,在大规模MIMIC-CXR数据集上的全面评估验证了模型的强泛化能力,可生成高质量、临床相关的报告。
原文摘要 · Abstract (English)
Medical imaging interpretation is a foundational pillar of modern clinical diagnostics, yet the manual generation of radiology reports remains a time-consuming process prone to interpretation inconsistencies. Within the field of medical AI, automating these descriptions through deep learning promises to streamline clinical workflows and standardise diagnostic output. However, accurate disease detection and precise report generation remain significant challenges due to limitations in capturing fine-grained visual features and ensuring clinical coherence. To address these issues, we propose RL-ACRGNet, an improved encoder-decoder model that integrates a pre-trained DenseNet encoder with a multilevel LSTM decoder within an off-policy reinforcement learning framework. Using a dual-network approach to refine visual-semantic embeddings through a metric-based reward mechanism, we demonstrate that RL-ACRGNet consistently outperforms state-of-the-art baselines on the IU-Xray dataset, achieving quantitative improvements in BLEU-4 (0.47%), METEOR (0.17%) and ROUGE-L (0.518). Furthermore, comprehensive evaluations on the large-scale MIMIC-CXR data set confirm the robust generalisation of the model and its ability to generate high-quality, clinically relevant reports
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。