用强化学习让多个智能体协同生成更准确的放射科报告。
Multi-Modal Multi-Agent Reinforcement Learning for Radiology Report Generation
- 将胸片解读拆分为局部与全局智能体,联合优化生成过程。
- 在MIMIC-CXR和IU数据集上提升报告临床有效性指标,达当前最优。
- 生成报告更准确、细节丰富,临床医生盲评媲美真实报告。
我们提出MARL-Rad,一种用于放射科报告生成的多模态多智能体强化学习框架,该框架在实际放射科工作流中训练整个智能体系统。针对现有方法中固定大模型被事后组织成手工设计的工作流而未针对角色优化的问题,MARL-Rad将胸片解读分解为区域特异性智能体与全局整合智能体,并使用可临床验证的奖励信号联合优化。在MIMIC-CXR和IU X-ray数据集上的实验表明,MARL-Rad在RadGraph、CheXbert和GREEN等临床有效性指标上持续提升,达到当前最优性能。进一步分析显示,该方法提升了左右侧一致性,生成报告更准确、更详细。盲评临床医生评估也表明,MARL-Rad生成的报告在临床上可与真实报告相媲美。
原文摘要 · Abstract (English)
We propose MARL-Rad, a multi-modal multi-agent reinforcement learning framework for radiology report generation that trains the entire agentic system on policy within its deployed radiology workflow. MARL-Rad addresses the limitation of post-hoc agentization, where fixed LLMs are organized into hand-designed agentic workflows without being optimized for their assigned roles. Our framework decomposes chest X-ray interpretation into region-specific agents and a global integrating agent, and jointly optimizes them using clinically verifiable rewards. Experiments on the MIMIC-CXR and IU X-ray datasets show that MARL-Rad consistently improves clinical efficacy metrics such as RadGraph, CheXbert, and GREEN scores, achieving state-of-the-art clinical efficacy performance. Further analyses show that MARL-Rad improves laterality consistency and produces more accurate and detailed reports. A blinded clinician evaluation further suggests that MARL-Rad produces reports clinically comparable to ground-truth reports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。