arXiv:2510.26498cs.CL2025-10

用多个大模型协作评估医疗AI分诊工具,效果比单个模型更可靠。

A Multi-agent Large Language Model Framework to Automatically Assess Performance of a Clinical AI Triage Tool

  • 用8个开源大模型+1个合规版GPT-4o组成集成系统,统一评估脑部CT报告
  • 最佳模型组合的MCC达0.571,显著高于单个模型的0.522
  • 适合医学AI评估、临床验证及模型可信度研究者参考

目的:评估多个大语言模型(LLM)协同是否比单一模型更可靠地评估基于像素的AI分诊工具性能。方法:使用商业颅内出血(ICH)检测工具处理来自14家医院的29,766例非对比头颅CT检查,由8个开源LLM模型与一个符合HIPAA的GPT-4o内部版本组成的集成系统,通过单一多示例提示评估报告中是否存在ICH。随机抽取1,726例进行人工复核,比较8个开源模型及共识结果与GPT-4o的性能表现。测试了三种理想化的模型集成方案对分诊工具的评分能力。结果:共纳入29,766对头颅CT图像-报告数据。Llama3.3:70b与GPT-4o分别取得最高AUC(0.78)和平均精度(0.75 & 0.76)。Llama3.3:70b在F1(0.81)与召回率(0.85)上最优,同时具备更高精确率(0.78)、特异度(0.72)与马修相关系数(MCC=0.57)。MCC 95%置信区间显示:全9模型集成组为0.571(0.552–0.591),前3模型集成组为0.558(0.537–0.579),共识组为0.556(0.539–0.574),而GPT-4o为0.522(0.500–0.543)。三类集成方式间无统计学差异(p > 0.05)。结论:中大型开源模型的集成系统,在回顾性评估临床AI分诊工具时,比单一模型更具一致性与可靠性。

原文摘要 · Abstract (English)

Purpose: The purpose of this study was to determine if an ensemble of multiple LLM agents could be used collectively to provide a more reliable assessment of a pixel-based AI triage tool than a single LLM. Methods: 29,766 non-contrast CT head exams from fourteen hospitals were processed by a commercial intracranial hemorrhage (ICH) AI detection tool. Radiology reports were analyzed by an ensemble of eight open-source LLM models and a HIPAA compliant internal version of GPT-4o using a single multi-shot prompt that assessed for presence of ICH. 1,726 examples were manually reviewed. Performance characteristics of the eight open-source models and consensus were compared to GPT-4o. Three ideal consensus LLM ensembles were tested for rating the performance of the triage tool. Results: The cohort consisted of 29,766 head CTs exam-report pairs. The highest AUC performance was achieved with llama3.3:70b and GPT-4o (AUC= 0.78). The average precision was highest for Llama3.3:70b and GPT-4o (AP=0.75 & 0.76). Llama3.3:70b had the highest F1 score (0.81) and recall (0.85), greater precision (0.78), specificity (0.72), and MCC (0.57). Using MCC (95% CI) the ideal combination of LLMs were: Full-9 Ensemble 0.571 (0.552-0.591), Top-3 Ensemble 0.558 (0.537-0.579), Consensus 0.556 (0.539-0.574), and GPT4o 0.522 (0.500-0.543). No statistically significant differences were observed between Top-3, Full-9, and Consensus (p > 0.05). Conclusion: An ensemble of medium to large sized open-source LLMs provides a more consistent and reliable method to derive a ground truth retrospective evaluation of a clinical AI triage tool over a single LLM alone.

AI医疗大模型集成医学影像评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。