通过优化视觉编码与压缩策略,提升3D医学影像报告生成效率与准确性。
Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation

- 采用解剖引导的感兴趣区域裁剪,有效减少冗余信息。
- 高分辨率输入配合PerceiverResampler投影器,使临床宏F1达49.5(CT-RATE)。
- 适用于追求高效高精度3D医学影像自动报告生成的研究者。
视觉-语言模型为自动化放射科报告生成提供了前景,但将其应用于完整的3D CT数据体时面临显著计算挑战。现代基础视觉编码器(VEs)每扫描可生成数万条视觉标记,导致传递给大语言模型(LLM)的视觉序列成为主要计算瓶颈。视觉到语言的投影器可压缩该序列以降低计算量,但可能丢失临床相关细节;而有效的压缩策略可在保持下游标记数量不变的前提下支持更高分辨率输入。如何在输入视野、空间分辨率及视觉-语言投影之间分配视觉标记预算仍是未解设计问题。我们系统评估了四种异构视觉编码器(基于CNN和ViT)、五种最高可达64倍压缩率的标记压缩投影器(含非压缩的MLP基线),以及五种指令微调的LLM(1.7B–4B)在两个大规模CT报告数据集(CT-RATE和Merlin)上的表现。在匹配的LLM标记预算下,解剖引导的感兴趣区域裁剪是最一致有效的策略,在20个实验设置中改善临床宏F1平均+3.7点(3D ViT Primus编码器)和+1.1点(2D ViT Curia编码器)。提升输入分辨率的效果高度依赖投影器:PerceiverResampler与更高分辨率的Curia特征结合,在两个数据集上均表现最优。最佳配置在测试集上达到当前最优临床宏F1,分别为49.5(CT-RATE)和49.0(Merlin)。代码与模型将在发表后公开。
原文摘要 · Abstract (English)
Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges. Modern foundation vision encoders (VEs) can produce tens of thousands of vision tokens per scan, making the visual sequence passed to the large language model (LLM) a primary computational bottleneck. Vision-to-language projectors can compress this sequence to reduce computation, but may discard clinically relevant detail; conversely, effective compression can accommodate higher-resolution inputs while keeping the downstream token count fixed. How this vision-token budget should be allocated across input field of view, spatial resolution, and vision-to-language projection therefore remains an open design question. We systematically evaluate four heterogeneous VEs (CNN- and ViT-based), five token-reducing projectors at up to 64x compression alongside a non-reducing MLP projector baseline, and five instruction-tuned LLMs (1.7B--4B) on two large-scale CT report datasets (CT-RATE and Merlin). At matched LLM token budgets, anatomy-guided region of interest cropping is the most consistent strategy, improving clinical macro F1 in 19 of 20 settings by +3.7 points on average for the 3D ViT Primus encoder and +1.1 for the slice-based 2D ViT Curia encoder. Increasing input resolution further is strongly projector-dependent: the PerceiverResampler, paired with higher-resolution Curia features, yields the strongest configuration in the resolution study on both datasets. Our best configurations achieve state-of-the-art clinical macro F1 on the test sets, reaching 49.5 on CT-RATE and 49.0 on Merlin. Code and models will be published upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。