arXiv:2511.20145cs.CV2025-11被引 3

用3D双分支模型自动生成淋巴瘤PET/CT报告,提升临床准确性。

Vision-Language Models for Automated 3D PET/CT Report Generation

  • 分通道编码PET与CT体积数据,结合风格自适应提示缓解医院间报告差异。
  • 在824份报告、24万张切片的数据集上,报告生成准确率提升31.49%。
  • 专为淋巴瘤设计评估体系,适合医学影像与AI交叉研究者使用。

正电子发射断层扫描/计算机断层扫描(PET/CT)在肿瘤学中至关重要,但扫描仪数量激增已超过专业医师供给,自动化报告生成(PETRG)成为减轻临床负担的关键。与结构成像不同,功能型PET受示踪剂生理影响大,需全身体积上下文理解。为此,本文提出端到端3D双分支框架PETRG-3D,分别编码PET与CT体积,并引入风格自适应提示以减少医院间报告习惯差异。构建了多中心淋巴瘤数据集PETRG-Lym(来自4家医院,824份报告,245,509对配对切片),并建立公开基准AutoPET-RG-Lym,基于开放影像数据生成135例专家撰写、临床验证的报告。为评估临床价值,提出淋巴瘤专用评估协议PETRG-Score,联合评估特定解剖区域内的代谢与结构发现。实验表明,PETRG-3D在自然语言指标(如ROUGE-L提升31.49%)和临床效能指标(如PET-All提升8.18%)上显著优于现有方法,验证了体积分模态建模与风格感知提示的优势。本工作为未来面向疾病认知推理与临床可信评估的PET/CT专用模型奠定基础。代码、模型及AutoPET-RG-Lym将公开。

原文摘要 · Abstract (English)

Positron emission tomography/computed tomography (PET/CT) is essential in oncology, yet the rapid expansion of scanners has outpaced the availability of trained specialists, making automated PET/CT report generation (PETRG) increasingly important for reducing clinical workload. Compared with structural imaging (e.g., X-ray, CT, and MRI), functional PET poses distinct challenges: metabolic patterns vary with tracer physiology, and whole-body 3D contextual information is required rather than local-region interpretation. To advance PETRG, we propose PETRG-3D, an end-to-end 3D dual-branch framework that separately encodes PET and CT volumes and incorporates style-adaptive prompts to mitigate inter-hospital variability in reporting practices. We construct PETRG-Lym, a multi-center lymphoma dataset collected from four hospitals (824 reports w/ 245,509 paired PET/CT slices), and construct AutoPET-RG-Lym, a publicly accessible PETRG benchmark derived from open imaging data but equipped with new expert-written, clinically validated reports (135 cases). To assess clinical utility, we introduce PETRG-Score, a lymphoma-specific evaluation protocol that jointly measures metabolic and structural findings across curated anatomical regions. Experiments show that PETRG-3D substantially outperforms existing methods on both natural language metrics (e.g., +31.49\% ROUGE-L) and clinical efficacy metrics (e.g., +8.18\% PET-All), highlighting the benefits of volumetric dual-modality modeling and style-aware prompting. Overall, this work establishes a foundation for future PET/CT-specific models emphasizing disease-aware reasoning and clinically reliable evaluation. Codes, models, and AutoPET-RG-Lym will be released.

医学影像生成模型多模态淋巴瘤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。