arXiv:2505.09787cs.AI2025-05被引 14

用多智能体模拟医生诊断流程,生成更准确可靠的影像报告

A Multimodal Multi-Agent Framework for Radiology Report Generation

  • 设计多个专业智能体协同完成检索、分析、撰写等步骤
  • 在自动评估和大模型评价中均优于基线,减少幻觉错误
  • 适合需要可解释性与高可信度的临床AI应用

放射科报告生成(RRG)旨在从医学影像自动生成诊断报告,有望提升临床效率并减轻放射科医生负担。尽管近期基于多模态大语言模型(MLLMs)和检索增强生成(RAG)的方法已取得显著进展,但仍面临事实不一致、幻觉及跨模态错位等问题。本文提出一种符合临床逐步推理流程的多模态多智能体框架,由任务专用智能体分别负责检索、初稿生成、视觉分析、优化与整合。实验结果表明,该方法在自动指标和基于大模型的评估中均优于强基线,生成的报告更具准确性、结构化与可解释性。本工作凸显了临床对齐的多智能体框架在支持可解释、可信临床AI应用方面的潜力。

原文摘要 · Abstract (English)

Radiology report generation (RRG) aims to automatically produce diagnostic reports from medical images, with the potential to enhance clinical workflows and reduce radiologists' workload. While recent approaches leveraging multimodal large language models (MLLMs) and retrieval-augmented generation (RAG) have achieved strong results, they continue to face challenges such as factual inconsistency, hallucination, and cross-modal misalignment. We propose a multimodal multi-agent framework for RRG that aligns with the stepwise clinical reasoning workflow, where task-specific agents handle retrieval, draft generation, visual analysis, refinement, and synthesis. Experimental results demonstrate that our approach outperforms a strong baseline in both automatic metrics and LLM-based evaluations, producing more accurate, structured, and interpretable reports. This work highlights the potential of clinically aligned multi-agent frameworks to support explainable and trustworthy clinical AI applications.

影像报告多智能体临床AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。