arXiv:2412.09529cs.CV2024-12被引 11

测试大模型在放射科任务中的代理核心能力,发现其处理复杂任务仍有瓶颈。

How Well Can Modern LLMs Act as Agent Cores in Radiology Environments?

  • 构建放射科评测平台,用2200个合成病例评估7个大模型表现
  • 复杂任务下模型整体完成率仅67.1%,工具协同能力弱是主要短板
  • 通过提示工程和多智能体协作,复杂任务成功率提升48.2%可参考

我们提出了RadA-BenchPlat评测平台,通过2,200个经放射科医生验证的合成患者记录(涵盖六个解剖部位、五种成像模态、2,200种疾病情景),生成24,200个问答对,模拟多样临床场景,用于评估大语言模型(LLMs)作为放射科代理核心的表现。平台定义了十类工具类别,并评估了七款主流大模型,结果显示,在常规任务中Claude-3.7-Sonnet可实现67.1%的任务完成率,但在复杂任务理解与工具协调方面仍存在明显不足,限制其作为自动化放射系统核心的能力。引入四种先进的提示工程策略后,复杂任务性能提升48.2%,其中提示反向传播和多智能体协作分别贡献16.8%和30.7%的改进。此外,探索自动化工具构建方法,实现65.4%的成功率,为未来全自动化放射应用落地提供可行路径。所有代码与数据已开源:https://github.com/MAGIC-AI4Med/RadABench。

原文摘要 · Abstract (English)

We introduce RadA-BenchPlat, an evaluation platform that benchmarks the performance of large language models (LLMs) act as agent cores in radiology environments using 2,200 radiologist-verified synthetic patient records covering six anatomical regions, five imaging modalities, and 2,200 disease scenarios, resulting in 24,200 question-answer pairs that simulate diverse clinical situations. The platform also defines ten categories of tools for agent-driven task solving and evaluates seven leading LLMs, revealing that while models like Claude-3.7-Sonnet can achieve a 67.1% task completion rate in routine settings, they still struggle with complex task understanding and tool coordination, limiting their capacity to serve as the central core of automated radiology systems. By incorporating four advanced prompt engineering strategies--where prompt-backpropagation and multi-agent collaboration contributed 16.8% and 30.7% improvements, respectively--the performance for complex tasks was enhanced by 48.2% overall. Furthermore, automated tool building was explored to improve robustness, achieving a 65.4% success rate, thereby offering promising insights for the future integration of fully automated radiology applications into clinical practice. All of our code and data are openly available at https://github.com/MAGIC-AI4Med/RadABench.

大模型放射科智能代理评测平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。