arXiv:2508.04062eess.IVcs.CV2025-08AAAI被引 6

首个面向PET影像的自动报告生成基准,提升肿瘤与神经疾病诊断效率

PET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography

  • 构建首个包含代谢信息的全身PET影像-报告配对数据集
  • 30个先进视觉语言模型在该任务上表现不佳,远未达临床实用水平
  • 引入临床有效性指标,评估放射性示踪剂摄取模式描述准确性

正电子发射断层扫描(PET)是现代肿瘤学和神经学影像的核心技术,能揭示传统解剖成像无法捕捉的动态代谢过程。放射科报告对临床决策至关重要,但人工撰写耗时费力。尽管视觉语言模型(VLMs)在医疗领域展现潜力,现有应用多集中于结构影像,忽视了分子PET成像的独特性。为此,我们提出PET2Rep,首个用于PET影像报告生成的大型综合性基准。该数据集首次涵盖全身图像-报告对,覆盖数十个器官,填补现有基准空白,更贴近真实临床需求。除通用自然语言生成指标外,我们引入一系列临床有效性指标,评估关键器官中放射性示踪剂摄取模式的描述质量。我们对30个前沿通用与医学专用VLM进行了横向对比,结果表明当前最优VLM在该任务上表现显著不足,难以满足实际应用需求,并识别出若干亟待解决的关键缺陷。

原文摘要 · Abstract (English)

Positron emission tomography (PET) is a cornerstone of modern oncologic and neurologic imaging, distinguished by its unique ability to illuminate dynamic metabolic processes that transcend the anatomical focus of traditional imaging technologies. Radiology reports are essential for clinical decision making, yet their manual creation is labor-intensive and time-consuming. Recent advancements of vision-language models (VLMs) have shown strong potential in medical applications, presenting a promising avenue for automating report generation. However, existing applications of VLMs in the medical domain have predominantly focused on structural imaging modalities, while the unique characteristics of molecular PET imaging have largely been overlooked. To bridge the gap, we introduce PET2Rep, a large-scale comprehensive benchmark for evaluation of general and medical VLMs for radiology report generation for PET images. PET2Rep stands out as the first dedicated dataset for PET report generation with metabolic information, uniquely capturing whole-body image-report pairs that cover dozens of organs to fill the critical gap in existing benchmarks and mirror real-world clinical comprehensiveness. In addition to widely recognized natural language generation metrics, we introduce a series of clinical efficacy metrics to evaluate the quality of radiotracer uptake pattern description in key organs in generated reports. We conduct a head-to-head comparison of 30 cutting-edge general-purpose and medical-specialized VLMs. The results show that the current state-of-the-art VLMs perform poorly on PET report generation task, falling considerably short of fulfilling practical needs. Moreover, we identify several key insufficiency that need to be addressed to advance the development in medical applications.

PET影像报告生成视觉语言模型医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。