arXiv:2502.00528cs.CVcs.CL2025-02被引 7

用弱监督方法构建PET/CT图文标注数据,训练出3D视觉定位模型。

Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings

  • 通过识别报告中的SUVmax和层号自动标注病灶位置。
  • 模型在FDG和DCFPyL检查中F1达0.78和0.75,整体达0.80。
  • 适合医学影像与NLP交叉研究者参考。

视觉语言模型可通过视觉定位将文本描述与图像中的具体位置关联,有望提升放射科报告质量。但该领域缺乏大型标注的图像-文本数据集,尤其在PET/CT中。本研究开发了一种自动化弱标签生成管道,从25,578例PET/CT检查中提取了11,356个句子-标签对,基于此训练了3D视觉语言模型ConTEXTual Net 3D。该模型结合大语言模型文本嵌入与3D nnU-Net,通过令牌级交叉注意力实现跨模态对齐。弱标注管道在251例中准确识别病灶位置达98%(246/251),仅7.5%需边界调整。ConTEXTual Net 3D F1得分为0.80,显著优于LLMSeg(F1=0.22)和2.5D版本(F1=0.53),但仍低于两名核医学医师(F1=0.94和0.91)。模型在FDG(F1=0.78)和DCFPyL(F1=0.75)检查中表现较好,但在DOTATE(F1=0.58)和Fluciclovine(F1=0.66)中下降。对不同大小病灶表现稳定,但低摄取病灶定位精度降低。研究证明该弱标注方法可有效构建标注数据集,推动3D视觉定位模型发展,但更大规模数据仍可能需用于缩小与医生性能差距。

原文摘要 · Abstract (English)

Vision-language models can connect the text description of an object to its specific location in an image through visual grounding. This has potential applications in enhanced radiology reporting. However, these models require large annotated image-text datasets, which are lacking for PET/CT. We developed an automated pipeline to generate weak labels linking PET/CT report descriptions to their image locations and used it to train a 3D vision-language visual grounding model. Our pipeline finds positive findings in PET/CT reports by identifying mentions of SUVmax and axial slice numbers. From 25,578 PET/CT exams, we extracted 11,356 sentence-label pairs. Using this data, we trained ConTEXTual Net 3D, which integrates text embeddings from a large language model with a 3D nnU-Net via token-level cross-attention. The model's performance was compared against LLMSeg, a 2.5D version of ConTEXTual Net, and two nuclear medicine physicians. The weak-labeling pipeline accurately identified lesion locations in 98% of cases (246/251), with 7.5% requiring boundary adjustments. ConTEXTual Net 3D achieved an F1 score of 0.80, outperforming LLMSeg (F1=0.22) and the 2.5D model (F1=0.53), though it underperformed both physicians (F1=0.94 and 0.91). The model achieved better performance on FDG (F1=0.78) and DCFPyL (F1=0.75) exams, while performance dropped on DOTATE (F1=0.58) and Fluciclovine (F1=0.66). The model performed consistently across lesion sizes but showed reduced accuracy on lesions with low uptake. Our novel weak labeling pipeline accurately produced an annotated dataset of PET/CT image-text pairs, facilitating the development of 3D visual grounding models. ConTEXTual Net 3D significantly outperformed other models but fell short of the performance of nuclear medicine physicians. Our study suggests that even larger datasets may be needed to close this performance gap.

视觉定位PET/CT多模态弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。