用多角色协作模拟放射科工作流程,提升CT报告生成的准确性和可靠性。
MARCH: Multi-Agent Radiology Clinical Hierarchy for CT Report Generation

- 分角色设计居民、住院医师和主任医生代理,模拟临床诊断流程。
- 在RadGenome-ChestCT上显著优于现有方法,临床真实度与语言准确率双高。
- 适合医疗AI开发人员及希望提升报告可信度的研究者参考。
自动化3D放射科报告生成常出现临床幻觉且缺乏人类诊疗中的迭代验证。尽管近期视觉-语言模型(VLMs)有所进展,但通常作为单一“黑箱”系统运行,缺乏临床工作流中的协作监督。为此,我们提出MARCH(多代理放射科临床层级框架),模仿放射科部门的专业层级结构,为不同代理分配专门角色。MARCH利用居民代理进行初始起草,并结合多尺度CT特征提取;多个住院医师代理通过检索增强进行修订;主任代理则协调迭代式、立场导向的共识讨论,以解决诊断分歧。在RadGenome-ChestCT数据集上,MARCH显著优于当前最优基线,在临床真实性与语言准确性方面均表现更优。本研究证明,建模类人组织结构能显著提升高风险医疗领域AI的可靠性。
原文摘要 · Abstract (English)
Automated 3D radiology report generation often suffers from clinical hallucinations and a lack of the iterative verification found in human practice. While recent Vision-Language Models (VLMs) have advanced the field, they typically operate as monolithic "black-box" systems without the collaborative oversight characteristic of clinical workflows. To address these challenges, we propose MARCH (Multi-Agent Radiology Clinical Hierarchy), a multi-agent framework that emulates the professional hierarchy of radiology departments and assigns specialized roles to distinct agents. MARCH utilizes a Resident Agent for initial drafting with multi-scale CT feature extraction, multiple Fellow Agents for retrieval-augmented revision, and an Attending Agent that orchestrates an iterative, stance-based consensus discourse to resolve diagnostic discrepancies. On the RadGenome-ChestCT dataset, MARCH significantly outperforms state-of-the-art baselines in both clinical fidelity and linguistic accuracy. Our work demonstrates that modeling human-like organizational structures enhances the reliability of AI in high-stakes medical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。