用AI专家团队自动选最佳降噪方案,提升低剂量PET图像质量。
VLM- and LLM-Driven Multi-Agent System for PET Image Denoising

- 构建视觉语言与大模型协同的多智能体系统,自主判断图像质量
- 在1/20和1/50低剂量下,PSNR和SSIM均优于UNet、GAN、DDPM
- 支持闭环反馈与回滚,适合临床影像科部署使用
正电子发射断层成像(PET)存在空间分辨率有限和信噪比低的问题,影响定量精度与病灶检出。基于深度学习的降噪方法虽有潜力,但实际应用中常需多个专用模型及专家干预,如识别运动伪影、估计噪声水平以选择合适去噪器,以及去噪后进行病灶量化评估。近期视觉语言模型(VLM)在图像质量理解与大语言模型(LLM)在上下文推理方面的进展,为自动化决策流程带来新机遇。受临床专家工作流程启发,我们提出一种基于VLM与LLM的多智能体PET去噪框架,可动态评估图像质量与病灶状态,自主选择最优去噪模型与参数,并支持闭环反馈与回滚机制。在Siemens Biograph Vision Quadra PET/CT数据上,对1/20和1/50低剂量条件下的图像进行了实验。模块独立评估验证了各组件可靠性,完整框架在两种剂量下均取得高于UNet、GAN和DDPM基线的PSNR与SSIM。初步结果表明,闭环多智能体框架能根据图像条件自适应调整去噪策略,具备可行性。
原文摘要 · Abstract (English)
Positron emission tomography (PET) imaging suffers from limited spatial resolution and low signal-to-noise ratio, which can compromise quantitative accuracy and lesion detectability. Deep learning-based denoising methods have demonstrated strong potential for improving PET image quality. However, their practical deployment in real-world settings remains challenging, often requiring multiple specialized models and expert interventions, such as identifying motion-induced misregistration artifacts, estimating noise levels to select an appropriate denoiser, and performing lesion-focused quantitative assessment after denoising. Recent advances in vision-language models (VLMs) for image quality understanding and large language models (LLMs) for contextual reasoning provide new opportunities for automated, decision-driven workflows. Inspired by expert workflows for PET image quality enhancement, we propose an VLM- and LLM-driven multi-agent PET denoising framework that dynamically assesses image quality and lesion status, autonomously selects optimal denoising models and parameters, and enables closed-loop feedback with rollback mechanisms. Experiments were conducted on Siemens Biograph Vision Quadra PET/CT data with 1/20 and 1/50 low-dose settings. Individual module evaluations demonstrated the reliability of the agentic components, while the complete framework achieved higher PSNR and SSIM than UNet, GAN, and DDPM baselines at both dose levels. These preliminary results demonstrate the feasibility of using a closed-loop multi-agent framework to adapt PET denoising strategies to different image conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。