用更少数据和更好策略,让AI生成更精准的放射科报告。
Rethinking the Efficiency and Effectiveness of Reinforcement Learning for Radiology Report Generation
- 按诊断多样性采样数据,减少训练样本量
- 通过诊断词加权优化,提升关键信息生成准确率
- 适合关注临床实用性的医疗AI研究者
放射科医生迫切需要全自动的AI报告生成系统,但现有方法临床实用性不足。强化学习(RL)有潜力解决这一问题,但在该任务中应用仍不充分。本文重新审视了强化学习在数据效率与优化效果方面对放射科报告生成(R2G)的作用。首先,分析数据量与质量的影响,发现质量比数量更重要;为此提出基于诊断多样性的数据采样策略,在减少样本数的同时保持性能。其次,观察到报告中多数词为模板式、无诊断价值,而关键术语频率低,易被优化忽略;为此提出诊断词加权策略优化(DiTPO),以诊断F1分数为奖励信号,通过规则或梯度机制显式建模不同词的重要性,优先生成临床关键内容。在MIMIC-CXR、IU-Xray和CheXpert Plus数据集上的实验表明,本框架达到当前最优性能,且所需强化学习训练样本显著减少。尤其在MIMIC-CXR上,仅用20%的样本即达到0.516的F1分数。
原文摘要 · Abstract (English)
Radiologists highly desire fully automated AI for radiology report generation (R2G), yet existing approaches fall short in clinical utility. Reinforcement learning (RL) holds potential to address these shortcomings, but its adoption in this task remains underexplored. In this paper, we revisit RL in terms of data efficiency and optimization effectiveness for R2G tasks. First, we explore the impact of data quantity and quality on the performance of RL in medical contexts, revealing that data quality plays a more critical role than quantity. To this end, we propose a diagnostic diversity-based data sampling strategy that enables comparable performance with fewer samples. Second, we observe that the majority of tokens in radiology reports are template-like and diagnostically uninformative, whereas the low frequency of clinically critical tokens heightens the risk of being overlooked during optimization. To tackle this, we introduce Diagnostic Token-weighted Policy Optimization (DiTPO), which directly optimizes for clinical accuracy by using a diagnostic F1 score as the reward signal. Unlike standard RL approaches that treat all tokens equally, DiTPO explicitly models the varying importance of different tokens through rule- or gradient-based mechanisms to prioritize clinically relevant content. Extensive experiments on the MIMIC-CXR, IU-Xray, and CheXpert Plus datasets demonstrate that our framework achieves state-of-the-art (SOTA) performance while requiring substantially fewer training samples in RL. Notably, on MIMIC-CXR, our framework attains an F1 score of 0.516 using only 20% of the RL training samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。