arXiv:2506.14540cs.LGcs.AI2025-06NeurIPS被引 5

为临床决策模型设计更实用的评估方法,兼顾校准性与实际成本差异。

Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs

  • 基于校准理论改进交叉熵,加入临床成本权重
  • 在不同患者比例下平均评估,提升模型鲁棒性
  • 适合关注真实医疗场景中误判代价的研究者

基于机器学习的临床决策支持系统广泛使用概率评分来辅助患者管理。然而,常用的评估指标如准确率和AUC-ROC无法充分反映临床核心需求:模型校准性、对分布偏移的鲁棒性以及对不对称错误成本的敏感性。本文提出一种兼具理论严谨性与实践可行性的评估框架,用于选择经校准的阈值分类器。该框架基于合适的评分规则(特别是Schervish表示),推导出一种调整后的交叉熵(对数评分),在临床相关范围内对不同类别平衡下的成本加权性能进行平均。所提方法简单易用,对临床部署条件敏感,能有效筛选出既校准良好又适应真实世界变化的模型。

原文摘要 · Abstract (English)

Machine learning-based decision support systems are increasingly deployed in clinical settings, where probabilistic scoring functions are used to inform and prioritize patient management decisions. However, widely used scoring rules, such as accuracy and AUC-ROC, fail to adequately reflect key clinical priorities, including calibration, robustness to distributional shifts, and sensitivity to asymmetric error costs. In this work, we propose a principled yet practical evaluation framework for selecting calibrated thresholded classifiers that explicitly accounts for the uncertainty in class prevalences and domain-specific cost asymmetries often found in clinical settings. Building on the theory of proper scoring rules, particularly the Schervish representation, we derive an adjusted variant of cross-entropy (log score) that averages cost-weighted performance over clinically relevant ranges of class balance. The resulting evaluation is simple to apply, sensitive to clinical deployment conditions, and designed to prioritize models that are both calibrated and robust to real-world variations.

临床AI模型评估校准性成本敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。