探索3D多模态大模型设计空间,提升CT报告生成效果
Exploring the Design Space of 3D MLLMs for CT Report Generation
- 系统测试不同视觉表示、投影器与大模型组合方案
- 融合分割掩码使报告质量提升,最高增益10%(GREEN分数)
- 发现模型大小不影响性能,且预训练体积不匹配反降低效果
多模态大语言模型(MLLM)为自动化放射科报告生成(RRG)提供了新路径。本文系统研究了用于3D CT报告生成的3D MLLM设计空间,涵盖视觉输入表征、投影器、大语言模型(LLMs)及微调技术。我们提出两种基于知识的报告增强方法,在AMOS-MM数据集1,687例上将GREEN评分提升最高达10%,并在MICCAI 2024 AMOS-MM挑战赛中取得第二名。结果表明,在相同训练协议下,报告生成性能与LLM规模无关;若原始视觉变换器(ViT)在较小体数据上预训练,更大体积输入未必带来性能提升。此外,结合分割掩码可显著改善生成效果。代码已公开于https://github.com/bowang-lab/AMOS-MM-Solution。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have emerged as a promising way to automate Radiology Report Generation (RRG). In this work, we systematically investigate the design space of 3D MLLMs, including visual input representation, projectors, Large Language Models (LLMs), and fine-tuning techniques for 3D CT report generation. We also introduce two knowledge-based report augmentation methods that improve performance on the GREEN score by up to 10%, achieving the 2nd place on the MICCAI 2024 AMOS-MM challenge. Our results on the 1,687 cases from the AMOS-MM dataset show that RRG is largely independent of the size of LLM under the same training protocol. We also show that larger volume size does not always improve performance if the original ViT was pre-trained on a smaller volume size. Lastly, we show that using a segmentation mask along with the CT volume improves performance. The code is publicly available at https://github.com/bowang-lab/AMOS-MM-Solution
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。