arXiv:2603.19275cs.CLcs.AI2026-03

通过中段训练提升医学影像报告自动生成效果

Improving Automatic Summarization of Radiology Reports through Mid-Training of Large Language Models

  • 在预训练后增加临床子领域中段训练,优化模型表现
  • 新模型在ROUGE-L和事实准确性上均超越传统方法
  • 适合需要小样本学习的医疗AI研发人员

放射科报告的自动摘要对减轻医生负担至关重要。以往研究多采用‘预训练-微调’策略适配大语言模型。本研究提出一种中段训练的子领域适应方法,探索三种策略:(1) 通用领域预训练,(2) 临床领域预训练,(3) 临床领域预训练后进行子领域中段训练。基于佛罗里达大学(UF)健康机构的大规模临床文本构建模型,并在OpenI和MIMIC-CXR等基准数据集上开展中段训练与微调实验。结果表明,中段训练模型GatorTronT5-Radio在文本指标(ROUGE-L)和事实性指标(RadGraph-F1)上均优于无中段训练的模型。该方法还展现出更强的少样本学习能力,缓解了先前研究中报道的‘冷启动’学习障碍。研究支持采用‘预训练-中段训练-微调’流程替代直接微调策略。

原文摘要 · Abstract (English)

Automatic summarization of radiology reports is an essential application to reduce the burden on physicians. Previous studies have widely used the "pre-training, fine-tuning" strategy to adapt large language models (LLMs) for summarization. This study proposed a subdomain adaptation through a mid-training method to improve summarization. We explored three adaptation strategies: (1) general-domain pre-training, (2) clinical-domain pre-training, and (3) clinical-domain pre-training followed by subdomain mid-training. We developed models using large-scale clinical text from the University of Florida (UF) Health and conducted mid-training and fine-tuning experiments using widely used benchmark datasets including OpenI and MIMIC-CXR. The experimental results show that the mid-trained model, GatorTronT5-Radio, achieved the best performance, outperforming models without mid-training in both text-based measures (ROUGE-L) and factuality measures (RadGraph-F1). Our mid-training methods also demonstrate better few-shot learning and could alleviate the "cold start" problem reported in previous studies as a learning barrier. Our findings support the use of "pre-training, mid-training, fine-tuning," instead of the widely used direct fine-tuning strategy.

医学生成大模型摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。