构建首个大模型生成PET/CT报告印象的基准与高效微调方案
PET-F2I: A Comprehensive Benchmark and Parameter-Efficient Fine-Tuning of LLMs for PET/CT Report Impression Generation
- 基于4.1万真实报告构建评测基准,提出临床导向评估指标
- 7B参数小模型在实体覆盖率上提升3倍,生成质量显著优于基线
- 适合医疗AI研发者、放射科医生及临床部署系统开发者参考
PET/CT成像在肿瘤学和核医学中至关重要,但将复杂影像发现转化为精准诊断结论耗时费力。尽管大语言模型在医学文本生成方面展现出潜力,其在高度专业化的PET/CT领域仍研究不足。本文提出PET-F2I-41K(PET Findings-to-Impression Benchmark),一个基于超过4.1万份真实报告的大规模基准数据集,用于评估大模型在生成PET/CT报告印象方面的表现。在该基准上,我们对27个模型进行了全面评测,涵盖前沿闭源模型、开源通用模型和医疗专用模型,并基于Qwen2.5-7B-Instruct通过LoRA方法开发出领域适配的7B模型(PET-F2I-7B)。除了标准自然语言生成指标(如BLEU-4、ROUGE-L、BERTScore),我们引入三个临床相关指标:实体覆盖率(ECR)、未覆盖实体率(UER)和事实一致性率(FCR),以评估诊断完整性和事实可靠性。实验表明,无论前沿模型还是医疗专用模型,在零样本设置下均表现不佳。相比之下,PET-F2I-7B在多项指标上实现显著提升(如0.708 BLEU-4),实体覆盖率较最强基线提升3.0倍,同时具备成本低、延迟小、隐私保护优等优势。本工作不仅推动了大模型在医学报告生成中的应用,也为构建可靠且可临床部署的PET/CT报告系统提供了标准化评估框架。
原文摘要 · Abstract (English)
PET/CT imaging is pivotal in oncology and nuclear medicine, yet summarizing complex findings into precise diagnostic impressions is labor-intensive. While LLMs have shown promise in medical text generation, their capability in the highly specialized domain of PET/CT remains underexplored. We introduce PET-F2I-41K (PET Findings-to-Impression Benchmark), a large-scale benchmark for PET/CT impression generation using LLMs, constructed from over 41k real-world reports. Using PET-F2I-41K, we conduct a comprehensive evaluation of 27 models across proprietary frontier LLMs, open-source generalist models, and medical-domain LLMs, and we develop a domain-adapted 7B model (PET-F2I-7B) fine-tuned from Qwen2.5-7B-Instruct via LoRA. Beyond standard NLG metrics (e.g., BLEU-4, ROUGE-L, BERTScore), we propose three clinically grounded metrics - Entity Coverage Rate (ECR), Uncovered Entity Rate (UER), and Factual Consistency Rate (FCR) - to assess diagnostic completeness and factual reliability. Experiments reveal that neither frontier nor medical-domain LLMs perform adequately in zero-shot settings. In contrast, PET-F2I-7B achieves substantial gains (e.g., 0.708 BLEU-4) and a 3.0x improvement in entity coverage over the strongest baseline, while offering advantages in cost, latency, and privacy. Beyond this modeling contribution, PET-F2I-41K establishes a standardized evaluation framework to accelerate the development of reliable and clinically deployable reporting systems for PET/CT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。