arXiv:2505.18530cs.MAcs.AI2025-05被引 9

用多智能体框架提升医学影像报告生成的准确性和全面性

MRGAgents: A Multi-Agent Framework for Improved Medical Report Generation with Med-LVLMs

  • 设计专用智能体分别处理不同疾病类别,针对性优化诊断描述
  • 在IU X光和MIMIC-CXR数据集上训练,显著减少正常结论的偏差
  • 适合医疗AI研究者与临床辅助系统开发者参考

医学大视觉语言模型(Med-LVLMs)已被广泛用于医学报告生成。尽管其性能已达先进水平,但仍存在将所有发现默认为正常的偏差,导致报告忽略关键异常,且常缺乏对放射学相关区域的完整描述。为此,我们提出医学报告生成智能体(MRGAgents),一种新型多智能体框架,通过微调针对不同疾病类别的专用智能体来解决此问题。基于IU X光和MIMIC-CXR数据集的子集进行训练,MRGAgents生成的报告在正常与异常发现之间实现更好平衡,并确保对临床相关区域的全面描述。实验表明,该方法优于当前最先进模型,在报告全面性和诊断实用性方面均有提升。

原文摘要 · Abstract (English)

Medical Large Vision-Language Models (Med-LVLMs) have been widely adopted for medical report generation. Despite Med-LVLMs producing state-of-the-art performance, they exhibit a bias toward predicting all findings as normal, leading to reports that overlook critical abnormalities. Furthermore, these models often fail to provide comprehensive descriptions of radiologically relevant regions necessary for accurate diagnosis. To address these challenges, we proposeMedical Report Generation Agents (MRGAgents), a novel multi-agent framework that fine-tunes specialized agents for different disease categories. By curating subsets of the IU X-ray and MIMIC-CXR datasets to train disease-specific agents, MRGAgents generates reports that more effectively balance normal and abnormal findings while ensuring a comprehensive description of clinically relevant regions. Our experiments demonstrate that MRGAgents outperformed the state-of-the-art, improving both report comprehensiveness and diagnostic utility.

医学报告多智能体视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。