arXiv:2412.06828cs.CLcs.AI2024-12被引 15

用多智能体协作提升放射科报告印象生成质量

Enhancing LLMs for Impression Generation in Radiology Reports through a Multi-Agent System

  • 分三步:检索相似报告、生成印象、审查反馈
  • 相比单模型,诊断准确率与表达清晰度显著提升
  • 适合医疗AI研发者和放射科辅助系统设计者

本研究提出名为RadCouncil的多智能体大语言模型框架,用于提升放射科报告中印象部分的生成质量。该框架包含三个专用智能体:1)检索智能体从向量数据库中查找相似报告;2)放射科医生智能体基于原始发现内容及检索到的范例生成印象;3)评审智能体对生成结果评估并提供反馈。以胸部X光为案例,通过BLEU、ROUGE、BERTScore等量化指标以及GPT-4的定性评估验证性能。实验显示,相较于单智能体方法,RadCouncil在诊断准确性、风格一致性与表达清晰度方面均有显著提升。研究证明,多个分工明确的LLM智能体协同可有效增强专业医疗任务的表现,推动更稳健、可扩展的医疗AI系统发展。

原文摘要 · Abstract (English)

This study introduces "RadCouncil," a multi-agent Large Language Model (LLM) framework designed to enhance the generation of impressions in radiology reports from the finding section. RadCouncil comprises three specialized agents: 1) a "Retrieval" Agent that identifies and retrieves similar reports from a vector database, 2) a "Radiologist" Agent that generates impressions based on the finding section of the given report plus the exemplar reports retrieved by the Retrieval Agent, and 3) a "Reviewer" Agent that evaluates the generated impressions and provides feedback. The performance of RadCouncil was evaluated using both quantitative metrics (BLEU, ROUGE, BERTScore) and qualitative criteria assessed by GPT-4, using chest X-ray as a case study. Experiment results show improvements in RadCouncil over the single-agent approach across multiple dimensions, including diagnostic accuracy, stylistic concordance, and clarity. This study highlights the potential of utilizing multiple interacting LLM agents, each with a dedicated task, to enhance performance in specialized medical tasks and the development of more robust and adaptable healthcare AI solutions.

多智能体放射科报告LLM应用医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。