通过结构化多模态时序嵌入,提升胸部X光报告生成的临床可用性。
CXRMate-2: Structured Multimodal Temporal Embeddings and Tractable Reinforcement Learning for Clinically Acceptable Chest X-ray Radiology Report Generation

- 用结构化时序嵌入和高分辨率特征压缩,实现可操作的强化学习优化。
- 在MIMIC-CXR上比MedGemma 1.5 (4B) 提升11.2% GREEN和24.4% RadGraph-XL得分。
- 生成报告45%被放射科医生认为可接受,阅读性更优但检出率仍不足。
胸部X光(CXR)报告生成模型在自动化指标上进展迅速,但其临床实用性仍存疑,因缺乏放射科医生的定性评估。本文提出CXRMate-2,一种先进的CXR报告生成模型,通过结构化多模态时序嵌入与高分辨率视觉特征压缩,实现可操作的强化学习(RL),以统一方式将视觉、文本和时间上下文纳入大语言模型解码器条件。该方法支持组相对策略优化(GRPO),使用新奖励函数提升生成报告与放射科医生报告的语义对齐。在MIMIC-CXR、CheXpert Plus和ReXgradient数据集上,CXRMate-2显著优于强基线,相较MedGemma 1.5 (4B) 在MIMIC-CXR上分别取得11.2%和24.4%的GREEN与RadGraph-XL提升。为直接对比生成报告与放射科医生报告,我们开展盲法随机回顾性评估:三位顾问级放射科医生在MIMIC-CXR测试集120个研究中比较生成报告与人工报告。生成报告在45%评分中被判定为可接受(偏好或等同于人工报告),且八种发现中有七种偏好无统计学差异。偏好人工报告主要源于更高召回率,而生成报告在可读性上始终更受青睐。这些结果明确了通往临床可接受的路径:提升召回率与细微病灶检测能力,是实现非劣于放射科医生报告的主要障碍,为辅助性、放射科主导的工作流中的前瞻性评估奠定基础。
原文摘要 · Abstract (English)
Chest X-ray (CXR) radiology report generation (RRG) models have shown rapid progress on automated metrics, yet their clinical utility remains uncertain due to limited qualitative evaluation by radiologists. We present CXRMate-2, a state-of-the-art CXR RRG model that enables tractable reinforcement learning (RL) through structured multimodal temporal embeddings and high-resolution visual feature compression, for efficient, unified conditioning of an LLM decoder on visual, textual, and temporal context from a study and its prior. This enables group relative policy optimisation (GRPO), where a proposed reward function is used to improve semantic alignment with radiologist reports. Across the MIMIC-CXR, CheXpert Plus, and ReXgradient datasets, CXRMate-2 achieves statistically significant improvements over strong benchmarks, including gains of 11.2% and 24.4% in GREEN and RadGraph-XL, respectively, on MIMIC-CXR relative to MedGemma 1.5 (4B). To directly compare CXRMate-2 against radiologist reporting, we conduct a blinded, randomised qualitative retrospective evaluation. Three consultant radiologists compare generated and radiologist reports across 120 studies from the MIMIC-CXR test set. Generated reports were deemed acceptable (defined as preferred or rated equally to radiologist reports) in 45% of ratings, with no statistically significant difference in preference rates for seven of the eight analysed findings. Preferences for radiologist reports were driven primarily by higher recall, while generated reports were consistently preferred for readability. Together, these results define a clear pathway to clinically acceptable CXR RRG. Improving recall and the detection of subtle findings represents the primary remaining barrier to non-inferiority with radiologist reporting, positioning CXR RRG for prospective evaluation in assistive, radiologist-led workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。