arXiv:2503.07032cs.CLcs.CV2025-03被引 2

用大模型提升医学影像质检效率,实测多个模型表现各异。

Multimodal Human-AI Synergy for Medical Imaging Quality Control: A Hybrid Intelligence Framework with Adaptive Dataset Curation and Closed-Loop Evaluation

  • 构建161张胸片+219份报告的标准化数据集,评估大模型质检能力。
  • Gemini 2.0-Flash在胸片任务中宏平均F1达90,DeepSeek-R1报告审计召回率62.23%领先。
  • 提出闭环评估框架,适合医疗AI开发者与影像科医生参考。

医学影像质量控制对准确诊断至关重要,但传统方法依赖人力且主观性强。本研究建立标准化数据集与评估框架,系统评估大语言模型(LLMs)在图像质量评估与报告标准化中的表现。首先构建并匿名化包含161张胸部X光片(CXR)和219份CT报告的数据集。随后基于召回率、精确率与F1分数,评估Gemini 2.0-Flash、GPT-4o与DeepSeek-R1等多款模型在检测技术错误与不一致方面的性能。实验表明,Gemini 2.0-Flash在胸片任务中达到90的宏平均F1,具备强泛化能力但细粒度表现有限;DeepSeek-R1在CT报告审核中召回率达62.23%,优于其他模型;其精简版本表现不佳,而InternLM2.5-7B-chat则展现出最高额外发现率,说明其覆盖范围广但精度较低。结果凸显大模型在医学影像质量控制中的潜力,尤其以DeepSeek-R1与Gemini 2.0-Flash表现突出。

原文摘要 · Abstract (English)

Medical imaging quality control (QC) is essential for accurate diagnosis, yet traditional QC methods remain labor-intensive and subjective. To address this challenge, in this study, we establish a standardized dataset and evaluation framework for medical imaging QC, systematically assessing large language models (LLMs) in image quality assessment and report standardization. Specifically, we first constructed and anonymized a dataset of 161 chest X-ray (CXR) radiographs and 219 CT reports for evaluation. Then, multiple LLMs, including Gemini 2.0-Flash, GPT-4o, and DeepSeek-R1, were evaluated based on recall, precision, and F1 score to detect technical errors and inconsistencies. Experimental results show that Gemini 2.0-Flash achieved a Macro F1 score of 90 in CXR tasks, demonstrating strong generalization but limited fine-grained performance. DeepSeek-R1 excelled in CT report auditing with a 62.23\% recall rate, outperforming other models. However, its distilled variants performed poorly, while InternLM2.5-7B-chat exhibited the highest additional discovery rate, indicating broader but less precise error detection. These findings highlight the potential of LLMs in medical imaging QC, with DeepSeek-R1 and Gemini 2.0-Flash demonstrating superior performance.

医学影像大模型质检多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。