arXiv:2511.01953q-bio.QMeess.IV2025-11被引 2

通过特征可分性判断病理图像分类中先验修正是否有效

Reliability Assessment Framework Based on Feature Separability for Pathological Cell Image Classification under Prior Bias

  • 用余弦相似度衡量特征空间的类内类间区分能力
  • 特征可分性对性能影响是先验偏移的412倍
  • 提供阈值指导何时该修正,适合医疗AI部署者

训练与部署数据间的先验概率偏移挑战基于深度学习的医学图像分类。标准修正方法通过重加权后验概率调整先验偏移,但效果不稳定。本文构建可靠性框架,识别在何种情况下先验修正能提升或损害性能。分析了303例结直肠癌标本的CD103/CD8免疫染色图像,共获得185,432张标注细胞图像,涵盖16种细胞类型。在1.1–20倍不同偏移比下训练ResNet模型。通过基于余弦相似度的似然质量评分量化特征可分性,反映学习特征空间中的类内与类间差异。采用多元线性回归、ANOVA和广义加性模型(GAMs)分析特征可分性、先验偏移、样本充分性与F1性能之间的关系。结果显示,特征可分性主导性能表现(β=1.650,p<0.001),其影响强度为先验偏移的412倍(β=0.004,p=0.018)。GAM分析显示高预测能力(R²=0.876),趋势多为线性。质量阈值0.294可有效识别需修正的情况(AUC=0.610)。得分>0.5的细胞类型无需修正即具鲁棒性,而得分<0.3的类型始终需要调整。结论:修正收益取决于特征提取质量而非偏移幅度。所提框架提供定量指导,实现选择性修正,促进高效可靠诊断AI部署。

原文摘要 · Abstract (English)

Background and objective: Prior probability shift between training and deployment datasets challenges deep learning-based medical image classification. Standard correction methods reweight posterior probabilities to adjust prior bias, yet their benefit is inconsistent. We developed a reliability framework identifying when prior correction helps or harms performance in pathological cell image analysis. Methods: We analyzed 303 colorectal cancer specimens with CD103/CD8 immunostaining, yielding 185,432 annotated cell images across 16 cell types. ResNet models were trained under varying bias ratios (1.1-20$\times$). Feature separability was quantified using cosine similarity-based likelihood quality scores, reflecting intra- versus inter-class distinctions in learned feature spaces. Multiple linear regression, ANOVA, and generalized additive models (GAMs) evaluated associations among feature separability, prior bias, sample adequacy, and F1 performance. Results: Feature separability dominated performance ($β= 1.650$, $p < 0.001$), showing 412-fold stronger impact than prior bias ($β= 0.004$, $p = 0.018$). GAM analysis showed strong predictive power ($R^2 = 0.876$) with mostly linear trends. A quality threshold of 0.294 effectively identified cases requiring correction (AUC = 0.610). Cell types scoring $>0.5$ were robust without correction, whereas those $<0.3$ consistently required adjustment. Conclusion: Feature extraction quality, not bias magnitude, governs correction benefit. The proposed framework provides quantitative guidance for selective correction, enabling efficient deployment and reliable diagnostic AI.

病理图像先验偏移可靠性评估特征可分性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。