arXiv:2505.00917stat.MEcs.AI2025-05ICML被引 8

多变量选择新方法,精准控制错误发现率

Multivariate Conformal Selection

  • 基于多变量非一致性得分构建置信区间,实现严格统计保证
  • 在模拟与真实数据上显著提升选中优质样本的能力
  • 适用于药物研发等多指标筛选场景,尤其适合高可靠性要求

从大规模数据集中筛选高质量候选样本在药物发现、精准医疗及大语言模型对齐中至关重要。传统置信选择(CS)虽能提供严格的不确定性量化,但仅适用于单变量响应和标量准则。为此,我们提出多变量置信选择(mCS),一种面向多变量响应场景的通用扩展。该方法引入区域单调性,并采用多变量非一致性得分构造置信p值,实现有限样本下的错误发现率(FDR)控制。我们提出两种变体:mCS-dist使用基于距离的得分,mCS-learn则通过可微优化学习最优得分。在模拟与真实数据集上的实验表明,mCS在保持FDR控制的同时显著提升了选择效能,为多变量选择任务提供了稳健框架。

原文摘要 · Abstract (English)

Selecting high-quality candidates from large datasets is critical in applications such as drug discovery, precision medicine, and alignment of large language models (LLMs). While Conformal Selection (CS) provides rigorous uncertainty quantification, it is limited to univariate responses and scalar criteria. To address this issue, we propose Multivariate Conformal Selection (mCS), a generalization of CS designed for multivariate response settings. Our method introduces regional monotonicity and employs multivariate nonconformity scores to construct conformal p-values, enabling finite-sample False Discovery Rate (FDR) control. We present two variants: mCS-dist, using distance-based scores, and mCS-learn, which learns optimal scores via differentiable optimization. Experiments on simulated and real-world datasets demonstrate that mCS significantly improves selection power while maintaining FDR control, establishing it as a robust framework for multivariate selection tasks.

多变量选择置信推断错误发现率统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。