arXiv:2609.05532cs.CVcs.LG2026-09

专用于头颈癌PET/CT诊断的AI模型,准确率显著优于通用模型。

A Specialized Large Multimodal Model for Interpreting PET/CT in Head and Neck Cancer

论文配图:A Specialized Large Multimodal Model for Interpreting PET/CT in Head and Neck Cancer
图 1 · 摘自论文原文
  • 分两阶段训练,先学基础影像特征,再聚焦肿瘤与淋巴结定位
  • 外部验证中肿瘤判别准确率达69%,淋巴结定位F1达0.66
  • 适合放射科医生快速辅助诊断和医学教学使用

头颈癌的PET/CT诊断因解剖复杂而具挑战性,催生了计算机辅助诊断需求。通用多模态大模型在医疗场景中受限于领域知识不足、隐私安全问题及输出冗长,因此亟需专用独立模型。本研究评估了基于大规模多中心数据集、定制训练流程与自回归训练的专用多模态大模型在头颈癌PET/CT自动解读中的可行性。采用两阶段课程学习:第一阶段使用28,000组图像-对话对(由两名放射科医生标注)学习模态类型与高代谢等基础信息;第二阶段使用12,975组对学习原发肿瘤存在性及颈部淋巴结转移的位置与是否存在。外部验证涵盖四家机构,设备多样。结果表明,该专用模型显著优于ChatGPT与原始LLaVA-NeXT,在第二阶段外部验证中,ROUGE-L、ROUGE-S、余弦相似度、精确率、召回率与F1分别达到0.8751、0.8794、0.8324、0.8794、0.8711、0.8751,而通用模型均低于0.1。内部肿瘤分类准确率为83.14±1.15%,外部为69.03±0.81%;淋巴结定位的相应指标为0.6389、0.6257、0.5287、0.5782、0.6371、0.6648。结论显示,专用多模态大模型在快速精准的PET/CT诊断支持与医学教育方面具有潜力,具备临床转化前景。

原文摘要 · Abstract (English)

Background: Diagnosing head and neck cancer using PET/CT is clinically challenging and time-consuming due to the anatomical complexity of the region, motivating computer-aided diagnosis (CAD). Generalist Large Multimodal Models (LMMs) remain limited in medical contexts by insufficient domain-specific knowledge, privacy and security concerns, and verbosity, motivating specialized standalone LMMs. Purpose: We evaluated the feasibility of a specialized LMM for automated PET/CT interpretation in head and neck cancer using a large-scale multi-institutional PET/CT dataset, a tailored training curriculum, and autoregressive training. Methods: LLaVA-NeXT was fine-tuned using a two-level curriculum with image-conversation pairs curated by two radiologists from public data. The dataset included clinically important annotations such as primary tumor presence and metastatic lymph node location. Level 1 used 28,000 image-conversation pairs to learn basic information, including modality type and hypermetabolism. Level 2 used 12,975 pairs to learn primary tumor presence and the existence and anatomical location of cervical lymph node metastases. External validation included four institutions with diverse imaging devices. Results: The specialized LMM substantially outperformed ChatGPT and LLaVA-NeXT. In Level-2 external validation, ROUGE-L, ROUGE-S, Cosine Similarity, Precision, Recall, and F1 were 0.8751, 0.8794, 0.8324, 0.8794, 0.8711, and 0.8751, while generalist models consistently scored below 0.1. Primary tumor classification accuracy was 83.14 +/- 1.15% internally and 69.03 +/- 0.81% externally. For lymph node localization, the corresponding scores were 0.6389, 0.6257, 0.5287, 0.5782, 0.6371, and 0.6648. Conclusion: Specialized LMMs show promising results for fast, accurate PET/CT-based diagnostic support and medical education, highlighting their potential for clinical translation.

医学影像多模态模型癌症诊断PET/CT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。