arXiv:2410.10220cs.CV2024-10中稿 · the "Workshop on I…

用扩散自编码器嵌入发现医学影像数据中的隐藏偏差与质量问题

Detecting Unforeseen Data Properties with Diffusion Autoencoder Embeddings using Spine MRI data

  • 通过扩散自编码器提取影像嵌入,捕捉数据内在特征
  • 能有效区分性别、年龄等受保护变量,识别出协议异常样本
  • 适合关注医疗数据质量与模型公平性的研究者使用

深度学习在医学影像中取得显著进展,依赖大规模数据提升诊断与预后能力。然而大规模数据常因受试者选择和采集过程引入固有误差。本文利用扩散自编码器(DAE)嵌入,揭示数据特征与偏见,包括性别等受保护变量的偏差及反映非预期协议变化的数据异常。基于11,186名德国国家队列(NAKO)参与者颈椎、胸椎和腰椎的矢状面T2加权磁共振图像进行实验。对比了风格生成网络(StyleGAN)与变分自编码器(VAE)。大规模数据评估表明,DAE嵌入能有效分离性别与年龄等受保护变量;通过t-SNE可视化,识别出头位定位差异等非预期协议变异;嵌入可定位使性别预测器失效的异常样本。结果表明,先进嵌入技术如DAE可用于检测医疗影像数据中的质量缺陷与偏见,提升深度学习模型在医疗应用中的可靠性与公平性,最终改善患者诊疗效果。

原文摘要 · Abstract (English)

Deep learning has made significant strides in medical imaging, leveraging the use of large datasets to improve diagnostics and prognostics. However, large datasets often come with inherent errors through subject selection and acquisition. In this paper, we investigate the use of Diffusion Autoencoder (DAE) embeddings for uncovering and understanding data characteristics and biases, including biases for protected variables like sex and data abnormalities indicative of unwanted protocol variations. We use sagittal T2-weighted magnetic resonance (MR) images of the neck, chest, and lumbar region from 11186 German National Cohort (NAKO) participants. We compare DAE embeddings with existing generative models like StyleGAN and Variational Autoencoder. Evaluations on a large-scale dataset consisting of sagittal T2-weighted MR images of three spine regions show that DAE embeddings effectively separate protected variables such as sex and age. Furthermore, we used t-SNE visualization to identify unwanted variations in imaging protocols, revealing differences in head positioning. Our embedding can identify samples where a sex predictor will have issues learning the correct sex. Our findings highlight the potential of using advanced embedding techniques like DAEs to detect data quality issues and biases in medical imaging datasets. Identifying such hidden relations can enhance the reliability and fairness of deep learning models in healthcare applications, ultimately improving patient care and outcomes.

医学影像数据偏差扩散模型嵌入分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。