用目标成员概率得分提升选择精度,尤其在复杂目标下表现更优。
Null-Calibrated Conformal Selection via Target-Membership Scores
- 采用目标成员概率作为评分标准,更契合选择任务本质。
- 在方差驱动的目标上,性能显著优于传统方法。
- 适合需要严格控制假发现率的稀有目标场景。
同调选择旨在识别测试样本中未知响应落在目标区域的个体,同时控制假发现率。现有方法多沿用预测导向的非一致性评分,如残差或截断残差。本文认为,选择任务的自然评分应为目标成员概率——它直接对应被选事件,且任何单调变换均可给出奈曼-皮尔逊最优排序。这一差异对均值单调目标无关紧要,但在区间型、方差驱动、多峰或多重条件目标中至关重要,因传统评分可能与选择能力错配。本文研究基于成员概率的同调选择,提出一种校准路径:空校准同调选择(NCCS),通过与已确认非目标校准样本比较测试得分。在零交换性假设下,NCCS可生成有限样本有效的零原假设p值,可与BY方法结合处理任意依赖,或与BH方法结合在标准正相关条件下使用。实验验证了该评分原则:在均值单调目标上成员概率评分与传统方法相当;在方差驱动目标上显著超越;在稀有目标情形下,经NCCS校准后,虽牺牲部分功效,但保障了有限样本下的零原假设有效性,避免直接经验FDP阈值法的反保守问题。
原文摘要 · Abstract (English)
Conformal selection aims to identify test candidates whose unknown responses fall in a target region while controlling the false discovery rate. Existing methods often inherit prediction-oriented nonconformity scores, such as residual or clipped residual scores, from conformal prediction. We argue that the natural score for selection is instead the target-membership probability. This score directly addresses the binary event being selected, and any monotone transform of it gives the Neyman--Pearson oracle ranking at a fixed null selection level. This distinction is irrelevant for mean-monotone targets, where conventional scores induce essentially the same ranking, but becomes important for interval-valued, variance-driven, multimodal, or multi-condition targets, where prediction-oriented scores can be misaligned with selection power. We study membership-score-based conformal selection and isolate one conformal calibration route, Null-Calibrated Conformal Selection (NCCS), which ranks test scores against confirmed non-target calibration examples. Under null exchangeability, NCCS yields finite-sample valid null p-values, which can be combined with BY under arbitrary dependence or with BH under standard positive-dependence conditions. Experiments support the score principle: membership scores match conventional scores on mean-monotone targets, substantially improve over mean-score selection on variance-driven targets, and, when calibrated by NCCS, trade power for finite-sample null validity in rare-target regimes where direct empirical-FDP thresholding can be anti-conservative.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。