首个3D PET报告模型,能精准定位病灶并生成专业描述。
PETAR: Localized Findings Generation with Mask-Aware Vision-Language Modeling for PET Automated Reporting
- 用3D掩码和焦点提示捕捉不足0.1%体积的病灶细节
- 在11,356个病灶上实现优于所有2D/3D基线的报告质量
- 医生实测验证有效,适合医学影像自动化研究者
生成3D正电子发射断层扫描(PET)的自动报告是医学影像中的重要挑战。由于全身体积数据复杂、临床发现多样且标注数据稀缺,自动化报告难以实现。为此,我们推出了PETARSeg-11K,首个大规模公开的病变级对应数据集,包含11,356个病灶描述与3D分割掩码。同时提出PETAR-4B,一种面向掩码感知的3D视觉语言模型,联合编码PET、CT及3D病灶分割掩码,采用3D焦点提示捕捉通常不足0.1%体积的微小病灶细节。自动评估显示,PETAR-4B显著优于所有2D与3D基线模型。首次针对自动化PET报告的人体评估中,五名医师参与验证了模型的临床实用性,并建立了自动化指标与专家判断间的相关性。本工作为3D医学视觉语言理解提供了基础数据集与新架构。
原文摘要 · Abstract (English)
Generating automated reports for 3D positron emission tomography (PET) is an important and challenging task in medical imaging. PET plays a vital role in oncology, but automating report generation is difficult due to the complexity of whole-body 3D volumes, the wide range of potential clinical findings, and the limited availability of annotated datasets. To address these challenges, we introduce PETARSeg-11K, the first large-scale, publicly available dataset that provides lesion-level correspondence between 3D PET/CT volumes and free-text radiological findings. It comprises 11,356 lesion descriptions paired with 3D segmentations. Second, we propose PETAR-4B, a 3D vision-language model designed for mask-aware, spatially grounded PET/CT reporting. PETAR-4B jointly encodes PET, CT, and 3D lesion segmentation masks, using a 3D focal prompt to capture fine-grained details of lesions that normally comprise less than 0.1% of the volume. Evaluations using automated metrics show PETAR-4B substantially outperforming all 2D and 3D baselines. A human study involving five physicians -- the first of its kind for automated PET reporting -- confirms the model's clinical utility and establishes correlations between automated metrics and expert judgment. This work provides a foundational dataset and a novel architecture, advancing 3D medical vision-language understanding in PET.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。