arXiv:2510.16057cs.CLcs.AI2025-10中稿 · IEEE BHI 2025

用双模型共识提升X光诊断可信度,准确率最高达91.3%。

Fusion-Augmented Large Language Models: Boosting Diagnostic Trustworthiness via Model Consensus

  • 基于输出相似性构建双大模型共识机制
  • 多模态输入下诊断准确率达91.3%(原单模型最高76.9%)
  • 适合希望提升AI诊断可信度的临床研究者

本研究提出一种新型多模型融合框架,利用ChatGPT与Claude两个先进大语言模型,在CheXpert数据集上提升胸部X光片解读的可靠性。从224,316张胸片中随机选取234例放射科医生标注样本,仅用图像提示时,ChatGPT和Claude的诊断准确率分别为62.8%和76.9%。采用95%输出相似性阈值的共识方法后,准确率提升至77.6%。为进一步评估多模态输入效果,基于MIMIC-CXR模板生成合成临床记录,对50例随机样本进行图像与文本联合评估,此时ChatGPT准确率达到84%,Claude为76%,共识准确率达91.3%。在两种实验条件下,基于一致性的融合策略始终优于单一模型。结果表明,融合互补模态并使用输出级共识可显著提升AI辅助放射诊断的可信度与临床价值,且计算开销极小。

原文摘要 · Abstract (English)

This study presents a novel multi-model fusion framework leveraging two state-of-the-art large language models (LLMs), ChatGPT and Claude, to enhance the reliability of chest X-ray interpretation on the CheXpert dataset. From the full CheXpert corpus of 224,316 chest radiographs, we randomly selected 234 radiologist-annotated studies to evaluate unimodal performance using image-only prompts. In this setting, ChatGPT and Claude achieved diagnostic accuracies of 62.8% and 76.9%, respectively. A similarity-based consensus approach, using a 95% output similarity threshold, improved accuracy to 77.6%. To assess the impact of multimodal inputs, we then generated synthetic clinical notes following the MIMIC-CXR template and evaluated a separate subset of 50 randomly selected cases paired with both images and synthetic text. On this multimodal cohort, performance improved to 84% for ChatGPT and 76% for Claude, while consensus accuracy reached 91.3%. Across both experimental conditions, agreement-based fusion consistently outperformed individual models. These findings highlight the utility of integrating complementary modalities and using output-level consensus to improve the trustworthiness and clinical utility of AI-assisted radiological diagnosis, offering a practical path to reduce diagnostic errors with minimal computational overhead.

医学AI多模态模型融合诊断可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。