用患者数据生成前列腺MRI,专家发现图像虽技术逼真却可能误导诊断。
Technically Plausible but Clinically Misleading? Expert Evaluation of Patient-Personalized Synthetic Prostate MRI
- 基于年龄和PSA值的3D扩散模型生成个性化MRI
- 专家评估显示合成图像临床可信度不一,存在误导风险
- 适合医学影像生成研究者及临床AI落地审慎评估者
磁共振成像(MRI)在前列腺癌评估中至关重要,但其获取成本高且耗时,成为临床路径中的主要瓶颈。个性化MRI合成有望通过生成针对个体患者的临床真实图像,减少对扫描成像的依赖。本文利用3D扩散模型,基于年龄和前列腺特异性抗原(PSA)等常规临床变量生成个性化前列腺MRI。在大型公开数据集上,合成图像在标准图像相似性指标下与扫描图像相当。随后,我们开展一项小规模单盲随机专家研究,选取具有不同临床特征的病例子集,由放射科医生通过配对评估和李克特量表评分,对比个性化与非条件合成图像的临床可信度、诊断信心及误导风险,并收集定性反馈。尽管个性化合成图像在技术上显得合理,但专家评估揭示了信心水平差异及嵌入临床数据后可能引发误导性线索的风险。结果表明,尽管个性化合成在技术上可行,但在实际临床流程中应用前仍需谨慎评估生成特征对医生解读的影响。
原文摘要 · Abstract (English)
Magnetic resonance imaging (MRI) is central to prostate cancer assessment, yet its acquisition is costly and time-consuming, making it a major bottleneck in the patient care pathway. Personalized MRI synthesis is a promising direction because it may allow the generation of clinically realistic images tailored to individual patients while reducing the dependence on scanner-acquired imaging. In this work, we in-vestigate whether patient-personalized MRI synthesis produces images that experts perceive as clinically plausible and how such personaliza-tion influences expert interpretation. We synthesize patient-personalized prostate MRI using a 3D diffusion model conditioned on routine pre-imaging clinical variables age and PSA. Using a large publicly available dataset, we first verify that synthesized images are technically compara-ble to scanner-acquired MRIs using standard image similarity metrics. We then conduct a pilot, single-blinded, randomized expert study on a subset of cases sampled to reflect variation in patient clinical pro-files. A radiologist compared personalized and unconditioned synthetic MRIs using paired assessments and Likert-scale ratings of clinical plausibility, confidence, and risk of misleading interpretation, supplemented by qualitative feedback. Although personalized synthetic MRIs appeared technically plausible, expert interpretation highlighted variability in con-fidence and potential risks of misleading cues when routine clinical data was embedded into image generation. These findings suggest that while personalized synthesis may be technically feasible, careful assessment is needed to understand how generated cues influence clinician interpreta-tion before such systems can be safely integrated into clinical workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。