用多智能体框架让AI生成并评估放射科报告,提升临床可信度。
Medical AI Consensus: A Multi-Agent Framework for Radiology Report Generation and Evaluation
- 设计十类专业智能体协同完成影像分析到报告生成全流程。
- 在公开数据集上实现报告质量与临床相关性双重提升。
- 适配LLM开发全周期,适合医疗AI研发与评估人员参考。
自动化放射科报告生成面临双重挑战:构建临床可靠的系统,以及设计严谨的评估协议。我们提出一种多智能体强化学习框架,作为多模态临床推理在放射科生态中的基准与评估环境。该框架将大语言模型(LLMs)和大视觉模型(LVMs)整合于由十个专业化智能体组成的模块化架构中,分别负责图像分析、特征提取、报告生成、审核与评估。此设计支持在智能体层面(如检测与分割准确率)和共识层面(如报告质量与临床相关性)进行细粒度评估。我们在公开放射科数据集上实现了基于chatGPT-4o的方案,其中LLMs作为评估者,并结合放射科医生反馈。通过将评估协议与LLM开发生命周期(预训练、微调、对齐、部署)相衔接,所提出的基准为可信赖的偏差驱动放射科报告生成提供了路径。
原文摘要 · Abstract (English)
Automating radiology report generation poses a dual challenge: building clinically reliable systems and designing rigorous evaluation protocols. We introduce a multi-agent reinforcement learning framework that serves as both a benchmark and evaluation environment for multimodal clinical reasoning in the radiology ecosystem. The proposed framework integrates large language models (LLMs) and large vision models (LVMs) within a modular architecture composed of ten specialized agents responsible for image analysis, feature extraction, report generation, review, and evaluation. This design enables fine-grained assessment at both the agent level (e.g., detection and segmentation accuracy) and the consensus level (e.g., report quality and clinical relevance). We demonstrate an implementation using chatGPT-4o on public radiology datasets, where LLMs act as evaluators alongside medical radiologist feedback. By aligning evaluation protocols with the LLM development lifecycle, including pretraining, finetuning, alignment, and deployment, the proposed benchmark establishes a path toward trustworthy deviance-based radiology report generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。