轻量级框架让小模型更准,大模型不改也能生成精准3DCT报告。
Revisiting LLM Adaptation for 3D CT Report Generation: A Study of Scaling and Diagnostic Priors

- 用诊断先验嵌入融合图像特征,仅训练少量参数。
- 1.6亿参数以下模型微调更有效,超10亿参数则冻结主干效果更好。
- 适合医疗文本生成场景,尤其数据少时避免幻觉和过拟合。
近年来,多模态学习(包括大语言模型与视觉-语言模型)在自然图像上表现出强适应性。但将其应用于医学领域,特别是三维(3D)影像时,面临计算复杂度高、体积分量依赖性强以及视觉特征与临床术语间的语义鸿沟等挑战。对有限医疗数据直接微调大模型易导致过拟合与临床幻觉,即语言流畅度优先于临床真实性。本研究系统考察了3D CT报告生成中的参数高效适配策略,提出RAD3D-Prefix——一种轻量级诊断先验条件化框架,通过整合图像嵌入与多标签诊断分类输出,保留关键临床信息并缩小语义差距。该模块保持大模型冻结,仅需极少可训练参数,有效降低小样本数据下的过拟合风险。在9610万至16亿参数的多个大模型上进行实验发现,小模型微调收益更高,而超过10亿参数的大模型仅训练轻量投影层即可实现性能、泛化性与效率的最优平衡。在多项自动指标与临床医生评估中,RAD3D-Prefix均优于同类参数高效基线,且在跨域数据上表现稳健,所需可训练参数远少于全量微调方案。
原文摘要 · Abstract (English)
Recent advances in multimodal learning, including large language models (LLMs) and vision-language models (VLMs), have demonstrated strong adaptability to natural images. However, extending their use to the medical domain, particularly for volumetric (3D) images, is challenging due to high computational complexity, volumetric dependencies and the semantic gap between visual features and clinical terminology. Naively fine-tuning LLMs on limited medical data often leads to overfitting and clinical hallucination, where linguistic fluency is prioritized over clinical factuality. In this study, we investigate parameter-efficient adaptation strategies for volumetric CT report generation and introduce RAD3D-Prefix, a lightweight diagnostic-prior conditioning framework that minimizes the need for extensive parameter training. This module integrates image embeddings with multi-label diagnostic classification logits, preserving critical clinical details while bridging the semantic gap. By keeping the LLM frozen, our method requires minimal trainable parameters and mitigates the risk of overfitting on small, domain-specific datasets. Through a systematic study spanning LLMs from 96.1M to 1.6B parameters, we find that fine-tuning is most beneficial for smaller LLMs, whereas freezing larger (~1B+ LLMs and training only lightweight projection layers provides a superior trade-off between performance, generalization, and computational efficiency. Across multiple automatic metrics and a clinical reader study, RAD3D-Prefix outperforms comparable parameter-efficient baselines and demonstrates strong out-of-domain generalization while using substantially fewer trainable parameters than fully fine-tuned alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。