arXiv:2608.25148cs.CVcs.AI2026-08中稿 · the HemaRAI 2026 w…

冷冻血液模型在真实场景中可靠性存疑,跨设备性能暴跌且预测不准确。

Can You Trust Frozen Hematology Foundation Models under Acquisition Shift?

论文配图:Can You Trust Frozen Hematology Foundation Models under Acquisition Shift?
图 1 · 摘自论文原文
  • 用线性探测和检索评估15个冻结模型在不同采集条件下的表现
  • 跨数据集准确率下降34%-72%,校准能力崩溃(ECE从0.004升至0.35)
  • 适合关注医疗影像部署可靠性的研究者与临床开发者

冷冻血液基础模型嵌入在本域白细胞识别上已达接近饱和的准确率(0.98-0.997),但临床应用需应对扫描仪、机构、染色及制备流程差异。我们对15个冻结编码器(血液、病理与通用视觉)在四个公开单细胞采集域上进行了审计,涵盖准确率鲁棒性与校准能力。本域线性探测宏F1已饱和,但跨数据集宏F1下降34%-72%,排名重排:以DinoBloom-L为最优,但在最偏移目标(MLL23)上降至第10名,落后于RedDino及多个通用与病理编码器。1-NN检索比源适配线性头更稳定(中位ρ 0.65 vs 0.45),但无一能普遍预测鲁棒性。校准性能崩溃:源训练探测器本域几乎准确(预期校准误差,ECE=0.004),离域时却高度自信出错(ECE=0.35),温度缩放迁移效果差。进一步审计发现MLL23是DinoBloom内部数据集;因唯一未暴露数据集即为源域,该基准无法分离暴露与设备偏移影响。无标签自适应与熵基选择在均衡评估下安全,但在真实类先验偏移下失效。类平衡再标准化(CBR)虽提升所有目标先验情景平均表现并部分改善校准,但仍存在编码器级异常与残余失准。因此,血液基础模型评估必须联合审计准确性、校准、暴露与类先验鲁棒性。

原文摘要 · Abstract (English)

Frozen hematology foundation-model (FM) embeddings reach near-saturated in-domain white-blood-cell (WBC) accuracy, but clinical deployment demands reliability across scanners, sites, stains and preparation pipelines. We audit 15 frozen encoders (hematology, pathology, and general vision) across four public single-cell acquisition domains along two axes: accuracy robustness and calibration. In-domain linear-probe macro-F1 is saturated (0.98-0.997), yet cross-dataset macro-F1 drops 34-72% and rankings re-order: DinoBloom-L, the in-domain best, falls to 10th of 15 on the most-shifted target (MLL23) at the benchmark's shared 224-px input, behind RedDino and several general and pathology encoders. Rank transfer is probe-dependent: 1-NN retrieval is more stable on average than a source-fitted linear head (median $ρ$ 0.65 vs 0.45), but neither probe universally predicts target robustness. Calibration also collapses: source-trained probes are nearly calibrated in-domain (expected calibration error, ECE, 0.004) but confidently wrong off-domain (ECE 0.35), and source-fitted temperature scaling transfers poorly. We further audit pretraining exposure and identify MLL23 as DinoBloom's internal cohort; because DinoBloom's only held-out dataset is also our source domain, this benchmark cannot isolate exposure from scanner-associated shift. Label-free adaptation and marginal-entropy-based model selection appear safe under balanced evaluation but fail under realistic WBC class-prior shift. Class-Balanced Re-standardization (CBR), a training-free pseudo-label-balanced feature normalization, improves all evaluated target-prior scenario means and partially improves calibration, although encoder-level exceptions and residual miscalibration remain. Hematology FM benchmarks must therefore jointly audit accuracy, calibration, exposure, and class-prior robustness.

基础模型医学影像鲁棒性校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。