医学影像模型选型不应只看准确率,还要综合考量可靠性多个维度。
CRS-Bench: A Reference-Relative Reliability Benchmark for Medical Image Encoders

- 构建多维度评测框架,从判别力、校准度等四方面评估模型可靠性
- 通过17,575次实验对比15个模型,发现传统指标无法完全反映真实表现
- 提出临床可靠性得分(CRS),帮助医生选择更稳定可靠的影像模型
预训练图像编码器在医学图像分类中至关重要,因专家标注成本高且任务特定数据集有限。随着模型从通用型向广义医学及专科专用型演进,选择合适表示成为关键建模决策。仅靠干净测试的判别能力(如AUROC)不足以支撑此选择:具有相似AUROC的编码器在校准性、标签效率及对采集扰动或分布偏移的鲁棒性上可能差异显著。本文提出CRS-Bench,一个面向多目标医学编码器选择的受控基准。该基准在皮肤科、眼科和放射科任务上评估15个预训练编码器家族,使用ISIC 2019、APTOS 2019和CheXpert数据集,并以CheXpert到MIMIC-CXR的机构迁移作为实际分布偏移场景,生成17,575条受控运行记录与3,515条种子聚合度量行。每个编码器在判别力、校准性、标签效率和鲁棒性四个操作可靠性维度上被刻画。我们通过临床可靠性得分(CRS)整合这些维度,该分数为帕累托感知、参考相对的评分,融合优势支配、性能均衡与最差轴表现。虽然AUROC与CRS正相关,但二者非等价:105对比较中有21对排序反转,平均绝对排名位移达1.87。配对种子自举分析揭示潘德玛、MedSigLIP和MedGemma构成一个稳定的领先可靠性梯队,而非统计上唯一的最优者。CRS-Bench提供了一种超越纯测试集AUROC的受控框架,用于从多维可靠性特征中选择医学图像编码器。
原文摘要 · Abstract (English)
Pretrained image encoders are central to medical image classification, where expert annotation is costly and task-specific cohorts are often limited. As the model space expands from general-purpose to broad-medical and specialty-specific encoders, selecting the representation becomes a substantive modeling decision. Clean-test discrimination alone is insufficient for this purpose: encoders with similar AUROC can differ in calibration, label efficiency, and stability under acquisition perturbations or distribution shift. We introduce CRS-Bench, a controlled benchmark for multi-objective medical encoder selection. CRS-Bench evaluates 15 pretrained encoder families across dermatology, ophthalmology, and radiology using ISIC 2019, APTOS 2019, and CheXpert, with CheXpert-to-MIMIC-CXR as an observed institutional shift, yielding 17,575 controlled run records and 3,515 seed-aggregated metric rows. Each encoder is characterized along four operational reliability dimensions: discrimination, calibration, label efficiency, and robustness. We summarize these dimensions using the Clinical Reliability Score (CRS), a Pareto-aware, reference-relative score combining dominance, profile balance, and worst-axis performance. AUROC and CRS are positively associated but not decision-equivalent: 21 of 105 pairwise orderings reverse, with a mean absolute rank displacement of 1.87. Paired-seed bootstrap analysis identifies PanDerm, MedSigLIP, and MedGemma as a stable leading reliability tier rather than a statistically resolved single leader. CRS-Bench provides a controlled framework for selecting medical image encoders from multi-axis reliability profiles rather than clean-test AUROC alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。