arXiv:2603.18050eess.IV2026-03被引 1

对比深度学习与手工特征模型在脑部MRI质量评估中的泛化能力

Quality assessment of brain structural MR images: Comparing generalization of deep learning versus hand-crafted feature-based machine learning methods to new sites

  • 用手工特征+传统机器学习和深度学习两种方法评估图像质量
  • 两者在新扫描仪上表现均不佳,但深度学习更擅长发现低质量图像
  • 深度学习计算快、无需复杂预处理,适合广泛部署

脑部结构MRI质量评估对大规模神经影像研究至关重要,运动伪影可能显著影响临床估计。尽管视觉评分仍是金标准,但耗时且主观。本研究评估了两种主流自动质量评估(AQA)方法:基于手工特征与传统机器学习的MRIQC,以及采用深度学习架构的CNNQC。使用来自17个不同站点的1,098例T1加权图像,通过留一站点外(LOSO)方法评估在已见站点与全新站点上的表现。结果表明,深度学习与传统机器学习方法均难以在新扫描仪或站点上良好泛化。尽管MRIQC在多数未见站点上准确率更高,但CNNQC在检测低质量扫描方面敏感性更强。由于深度学习方法如CNNQC具有更高的计算效率且无需昂贵预处理,未来若能提升跨站点泛化能力,将更适合大规模部署。

原文摘要 · Abstract (English)

Quality assessment of brain structural MR images is critical for large-scale neuroimaging studies, where motion artifacts can significantly bias clinical estimates. While visual rating remains the gold standard, it is time-consuming and subjective. This study evaluates the relative performance and generalization capabilities of two prominent Automated Quality Assessment (AQA) methods: MRIQC, which uses hand-crafted image-quality metrics with traditional machine learning, and CNNQC, which utilizes a deep learning (DL) architecture. Using a heterogeneous dataset of 1,098 T1-weighted volumes from 17 different sites, we assessed performance on both seen sites and entirely new sites using a leave-one-site-out (LOSO) approach. Our results indicate that both DL and traditional ML methods struggle to generalize to new scanners or sites. While MRIQC generally achieved higher accuracy across most unseen sites, CNNQC demonstrated higher sensitivity for detecting poor-quality scans. Given that DL-based methods like CNNQC offer higher computational efficiency and do not require expensive pre-processing, they may be preferred for widespread deployment, provided that future work focuses on improving cross-site generalizability.

医学影像质量评估深度学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。