多模态大模型能用影像补全文本缺失,提升胸片报告准确性。
Utility of Multimodal Large Language Models in Analyzing Chest X-ray with Incomplete Contextual Information
- 融合图像与文本的多模态模型更抗数据缺失
- 图像输入使模型在文本缺20%-80%时性能仍稳定
- 适合临床场景中报告不完整时的辅助诊断
背景:大语言模型(LLM)在临床中应用日益广泛,但其性能常受放射科报告不完整影响。本研究测试了多模态大语言模型(结合文本与图像)在胸部放射学报告中的表现,以提升临床决策支持的准确性和可解释性。方法:使用MIMIC-CXR数据库中的300对影像-报告数据,评估OpenFlamingo、MedFlamingo和IDEFICS三个模型在纯文本和多模态模式下的表现。先基于完整文本生成印象,再逐步删除20%、50%、80%的文本内容。通过添加胸片影像,评估模型性能,采用三种指标进行统计分析。结果:纯文本模型表现相近(ROUGE-L: 0.39 vs. 0.21 vs. 0.21;F1RadGraph: 0.34 vs. 0.17 vs. 0.17;F1CheXbert: 0.53 vs. 0.40 vs. 0.40),OpenFlamingo在完整文本下最优(p<0.001)。所有模型在文本缺失时性能下降。但加入图像后,MedFlamingo和IDEFICS性能显著提升(p<0.001),在文本缺失情况下仍可达到或超过OpenFlamingo水平。结论:单靠文本的模型在数据不完整时输出质量低,而多模态模型可通过图像信息增强可靠性,助力临床决策。
原文摘要 · Abstract (English)
Background: Large language models (LLMs) are gaining use in clinical settings, but their performance can suffer with incomplete radiology reports. We tested whether multimodal LLMs (using text and images) could improve accuracy and understanding in chest radiography reports, making them more effective for clinical decision support. Purpose: To assess the robustness of LLMs in generating accurate impressions from chest radiography reports using both incomplete data and multimodal data. Material and Methods: We used 300 radiology image-report pairs from the MIMIC-CXR database. Three LLMs (OpenFlamingo, MedFlamingo, IDEFICS) were tested in both text-only and multimodal formats. Impressions were first generated from the full text, then tested by removing 20%, 50%, and 80% of the text. The impact of adding images was evaluated using chest x-rays, and model performance was compared using three metrics with statistical analysis. Results: The text-only models (OpenFlamingo, MedFlamingo, IDEFICS) had similar performance (ROUGE-L: 0.39 vs. 0.21 vs. 0.21; F1RadGraph: 0.34 vs. 0.17 vs. 0.17; F1CheXbert: 0.53 vs. 0.40 vs. 0.40), with OpenFlamingo performing best on complete text (p<0.001). Performance declined with incomplete data across all models. However, adding images significantly boosted the performance of MedFlamingo and IDEFICS (p<0.001), equaling or surpassing OpenFlamingo, even with incomplete text. Conclusion: LLMs may produce low-quality outputs with incomplete radiology data, but multimodal LLMs can improve reliability and support clinical decision-making. Keywords: Large language model; multimodal; semantic analysis; Chest Radiography; Clinical Decision Support;
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。