arXiv:2608.28714eess.IVcs.CV2026-08

6.8%的脑部MRI重建研究同时评估了图像保真度和医生读片,暴露了当前评价体系的重大盲区。

Evaluating the Safety of Deep Learning-Based Brain MRI Reconstruction

  • 通过系统综述整合263项研究,分析现有评估方法
  • 仅6.8%的研究同步进行图像质量与医生读片评估
  • 生成模型幻觉问题严重但极少被医生验证

目的:深度学习使脑部MRI加速4至10倍,但模型可能抹除病灶或合成虚假组织——这些错误在像素级指标如PSNR和SSIM下难以察觉。我们评估当前评价实践是否能发现这一盲点。方法:遵循PRISMA 2020,无时间限制地检索七个数据库,共纳入263项研究(1995–2026),使用QUADAS-2评估质量并匹配工具,进行叙述性综合。分类基于标题、摘要及受控词汇;报告的流行率代表下限。重复筛选达成高一致性(Fleiss kappa = 0.877),评估一致性为0.788(可观测部分为0.390)。数据提取未经审计。结果:仅18项研究(6.8%)在同一数据上同时记录了保真度指标和医生读片评估,核心替代指标未被测量。医生研究多关注组间一致性,结果较弱:fastMRI 2020中一致性分别为0.457和0.386(Kendall W),仅在SSIM偏离时改善。在所述误差模型下,抹除100 mm³腔隙性梗死仅使全局PSNR变化0.03 dB。随着文献量增长五倍,医生评估比例从32%降至18%,后回升至21%。生成模型(最常关联幻觉,占比39%)是医生评估最少的(11.3%),而自监督模型达47%且零医生评估。仅5%发布代码并开展医生评估;无人评估模型观察者;无命名数据集涵盖急性卒中或出血。结论:基于这些下限,当前评价实践无法保证诊断安全性。我们提出五项面向安全的评估必须满足的要求。

原文摘要 · Abstract (English)

Objective: Deep learning accelerates brain MRI four- to tenfold, but models can erase lesions or synthesize false tissue - failures pixel-averaged metrics like PSNR and SSIM miss. We review whether current evaluation practices detect this blind spot. Methods: Following PRISMA 2020, we searched seven databases without date limits, including 263 studies (1995-2026), appraised them using QUADAS-2 and matched instruments, and synthesized narratively. Categories were derived from titles, abstracts, and controlled vocabulary; reported prevalence figures represent floors. Duplicate screening achieved high agreement (Fleiss kappa = 0.877), as did appraisal (0.788; 0.390 where observable). Extraction is unaudited. Results: Only 18 of 263 studies (6.8%) recorded both a fidelity metric and reader assessment on identical data, leaving the central surrogate unmeasured. Reader studies mostly measured inter-reader agreement, which was weak: fastMRI 2020 concordance reached 0.457 and 0.386 (Kendall W), improving only where SSIM diverged. Erasing a 100 mm3 lacunar infarct shifts global PSNR by 0.03 dB under the stated error model. As the corpus grew fivefold, reader assessments dropped from 32% to 18%, recovering to 21%. Generative models - most associated with hallucination (39%) - were among the least reader-evaluated (11.3%), while self-supervised models reached 47% with zero reader evaluation. Only 5% released code and ran reader studies; none evaluated a model observer; no named dataset covered acute stroke or hemorrhage. Conclusions: On these floors, current evaluation practices cannot certify diagnostic safety. We derive five requirements safety-oriented evaluations must meet.

医学影像深度学习安全评估方法MRI重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。