用视觉问答模型生成3D胸部CT报告,提升诊断准确性和报告质量。
3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models
- 基于VQA框架构建3D CT报告生成模型
- 在公开与私有数据集上显著优于现有方法
- 适合临床放射科报告自动化需求
医学影像分析在现代放射学诊断中至关重要,尤其面对医学影像数据的指数级增长。自动化报告生成系统的需求日益迫切。以往研究多集中于2D医学影像的机器学习与多模态语言模型,而3D医学影像报告生成因数据稀缺和计算复杂性较少被探索。本文提出3D-CT-GPT,一种专为生成3D CT扫描(尤其是胸部CT)放射科报告设计的视觉问答(VQA)型医学视觉语言模型。在公开与私有数据集上的大量实验表明,3D-CT-GPT在报告准确性和质量方面显著优于现有方法。尽管当前方法有限,包括部分开源的CT2Rep和开源的M3D,我们通过适当的转换与评估方法确保公平比较。实验结果表明,3D-CT-GPT提升了诊断准确性与报告连贯性,确立了其作为临床放射科报告生成的可靠解决方案。未来工作将聚焦于数据集扩展与模型优化,以进一步提升性能与适用性。
原文摘要 · Abstract (English)
Medical image analysis is crucial in modern radiological diagnostics, especially given the exponential growth in medical imaging data. The demand for automated report generation systems has become increasingly urgent. While prior research has mainly focused on using machine learning and multimodal language models for 2D medical images, the generation of reports for 3D medical images has been less explored due to data scarcity and computational complexities. This paper introduces 3D-CT-GPT, a Visual Question Answering (VQA)-based medical visual language model specifically designed for generating radiology reports from 3D CT scans, particularly chest CTs. Extensive experiments on both public and private datasets demonstrate that 3D-CT-GPT significantly outperforms existing methods in terms of report accuracy and quality. Although current methods are few, including the partially open-source CT2Rep and the open-source M3D, we ensured fair comparison through appropriate data conversion and evaluation methodologies. Experimental results indicate that 3D-CT-GPT enhances diagnostic accuracy and report coherence, establishing itself as a robust solution for clinical radiology report generation. Future work will focus on expanding the dataset and further optimizing the model to enhance its performance and applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。