arXiv:2605.02942cs.LGcs.CV2026-05

揭示胎儿超声中时间与采集偏差的交叉影响,避免误判为性别等人口特征问题。

Intersectional Disentangling of Temporal and Acquisition Bias in Fetal Ultrasound

论文配图:Intersectional Disentangling of Temporal and Acquisition Bias in Fetal Ultrasound
图 1 · 摘自论文原文
  • 通过无监督切片发现,误差高的子群体共享极端的图像采集间距和扫描至分娩间隔。
  • 固定扫描-分娩间隔后,采集间距的影响几乎消失(β=-0.17),而时间间隔主导误差(β=+0.56)。
  • 深度学习模型比临床公式更敏感于时间偏差,但整体仍更准确,提示需区分真实性能与潜在混淆因素。

医学影像AI的公平性研究常将子群体表现差异归因于训练数据中的代表性不足。我们表明,交叉分析可解耦由临床和采集混杂因素引发的公平性与性能差距,这些因素与目标变量共变。以产科超声中扫描时间胎儿体重估计为例,分析了两种模型:最先进的深度学习(DL)模型和临床金标准Hadlock公式。通过无监督切片发现,高误差子群体在图像采集像素间距(PS)和扫描至分娩间隔(STD)上均呈现极端值。其中,PS是可优化的采集参数,而STD是PS及偏差诊断的潜在混杂因子。仅依靠子群体检查无法分离二者。采用模型无关分析并固定元数据分区与部分回归,发现固定STD时,PS效应降至微小残差(标准化系数β = -0.17);而固定PS时,STD主导误差(β = +0.56)。两种模型随STD增加而性能下降,包括生物测量公式,表明大部分误差源于预测目标本身的固有特性,而非成像。深度学习模型对STD的敏感度约为Hadlock公式的1.5倍,尽管如此,其在每个子群体中仍更准确。结论是,公平性分析必须仔细审视潜在混杂因素,否则可能将时间或采集相关依赖错误归因于人口统计学特征。

原文摘要 · Abstract (English)

Fairness studies of medical imaging AI often explain subgroup performance gaps through under-representation in the training data. We show that intersectional analysis can disentangle fairness and performance gaps arising from clinical and acquisition confounders that co-vary with the target. As a case, we study scan-time fetal weight estimation from obstetric ultrasound, analyzing two models: a state-of-the-art deep learning (DL) model and the clinical gold-standard Hadlock formula. Using unsupervised slice discovery, we find that high-error subgroups share extreme in the image-acquisition pixel spacing (PS) and in the scan-to-delivery (STD) interval. Of these, PS is an acquisition parameter that can be optimized, while STD is a potential confounder for both PS and our bias diagnostics. Subgroup inspection alone cannot separate them. We disentangle the factors using a model-agnostic analysis with identical metadata partitions and partial regression. Holding STD fixed, the apparent PS effect collapses to a small residual (standardized coefficient $β=-0.17$), whereas holding PS fixed, STD dominates error ($β=+0.56$). Both models degrade with increasing STD, including the biometric formula, indicating much of the error is intrinsic to the prediction target rather than imaging. The DL model is $\sim$1.5$\times$ more sensitive to STD than Hadlock, though it remains more accurate in every subgroup. We conclude that fairness analyses need to carefully analyze potential confounds, or risk attributing an effect such as temporal or acquisition-related dependency to demographics.

医学影像公平性超声混杂因素

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。