用强化学习提升医学影像报告生成,效果超越传统方法。
Scaling medical imaging report generation with multimodal reinforcement learning
- 采用强化学习直接优化临床评估指标,避免模板化问题。
- 在ReXrank基准上达到新SOTA,显著优于现有方法。
- 适用于多机构、多场景的医学影像报告生成任务。
前沿模型在自然语言理解与推理方面表现出色,但在生物医学等高价值垂直领域仍存在多模态理解与推理能力短板。医学影像报告生成是典型挑战。监督微调虽能提升性能,但易过拟合于表面模板模式。本文提出通用医学影像报告生成框架UniRG,通过强化学习作为统一机制,直接优化面向实际应用的评估指标,显著超越监督微调,并在不同医疗机构和临床实践中实现持久泛化。我们在公开胸片(CXR)数据上训练了UniRG-CXR,开展了严格的多场景评估。在权威的ReXrank基准上,UniRG-CXR取得新SOTA,大幅领先于先前最优方法。模型已开源:https://huggingface.co/microsoft/UniRG-CXR。
原文摘要 · Abstract (English)
Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit major competency gaps in multimodal understanding and reasoning especially in high-value verticals such as biomedicine. Medical imaging report generation is a prominent example. Supervised fine-tuning can substantially improve performance, but they are prone to overfitting to superficial boilerplate patterns. In this paper, we introduce Universal Report Generation (UniRG) as a general framework for medical imaging report generation. By leveraging reinforcement learning as a unifying mechanism to directly optimize for evaluation metrics designed for end applications, UniRG can significantly improve upon supervised fine-tuning and attain durable generalization across diverse institutions and clinical practices. We trained UniRG-CXR on publicly available chest X-ray (CXR) data and conducted a thorough evaluation in CXR report generation with rigorous evaluation scenarios. On the authoritative ReXrank benchmark, UniRG-CXR sets new overall SOTA, outperforming prior state of the art by a wide margin. We release our model at https://huggingface.co/microsoft/UniRG-CXR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。