为顺序搜索设计新评分规则,更准确评估模型诊断成本。
Pandora's Regret: A Proper Scoring Rule for Evaluating Sequential Search

- 基于搜索成本构建配对加性评分规则,捕捉排名关系
- 在597个MedMNIST模型上显著优于传统指标
- 适合医疗诊断等需考虑排序与成本的场景
在顺序搜索中,模型逐个测试候选项直至找到正确类别。标准评分规则如对数损失是局部的,忽略竞争项的排序,导致评估与实际搜索效用不一致。我们发现顺序搜索诱导出配对结构,并通过分析不同测试成本下的最优搜索期望成本,推导出潘多拉遗憾:一种闭式、配对可加且严格合理的评分规则。该规则既能激发真实概率,又能惩罚使干扰项排在真类之前的排序错误。其构造产生一个单参数贝塔族,可在排序混淆惩罚与概率大小间权衡,同时保持作为期望搜索成本的明确解释。我们证明了对数损失、准确率和宏F1依赖于与顺序搜索不匹配的隐含决策模型。在597个MedMNIST模型上的实验表明,基于潘多拉遗憾的指标比标准方法更准确预测临床诊断成本,将决策论评分规则推广至多分类场景。
原文摘要 · Abstract (English)
In sequential search, alternatives are tested until the true class is found. Standard proper scoring rules like log loss are local, ignoring the ranking of competitors and misaligning model evaluation with search utility. We show that sequential search induces a pairwise structure that overcomes this. By analyzing the expected cost of optimal search under varying testing costs, we derive Pandora's Regret: a closed-form, pairwise-additive, and strictly proper scoring rule. Pandora's Regret both elicits true probabilities and penalizes rank-reversing miscalibrations where distractors outrank the true class. Our construction yields a one-parameter Beta family that balances penalties for rank-swapping versus probability magnitude, while retaining a grounded interpretation as expected search cost. We prove that log loss, accuracy, and macro-F1 rely on implicit decision models misaligned with sequential search. Across 597 MedMNIST models, Pandora-based metrics better predict clinical diagnostic costs than standard alternatives, extending decision-theoretic scoring rule construction to the multiclass setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。