让大模型输出更可信,精准控制错误率并提升可用率
LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
- 用线性期望约束重构选择决策,直接控制被接受样本的错误概率
- 在有限样本下实现风险可控,保留率显著高于基线方法
- 适用于问答和视觉问答系统,适合对可靠性要求高的场景
基础模型常生成不可靠答案,而启发式置信度估计无法充分区分正确与错误输出,导致用户在无统计保证下接受错误结果。本文提出LEC框架,通过选择条件风险控制,确保被接受预测的错误概率不超过用户指定的风险水平。将选择性预测重新建模为受线性期望约束的决策问题,直接控制被接受错误数与被接受预测数的期望比,对应于选择后的边际错误概率。在可交换性假设下,仅需一个独立校准集即可推导出有限样本下的充分条件,从而计算出风险受限且保留率最大的阈值。进一步将LEC扩展至双模型路由系统:当主模型置信度超过校准阈值时,输入将转至次级模型,同时保持系统级的选择条件误差控制。在封闭式与开放式问答及视觉问答任务上的实验表明,LEC能有效维持预设风险水平,并显著提升样本保留率。
原文摘要 · Abstract (English)
Foundation models often generate unreliable answers, while heuristic uncertainty estimators fail to fully distinguish correct from incorrect outputs, causing users to accept erroneous answers without any statistical guarantee. We address this problem through selection-conditioned risk control, aiming to ensure that an accepted prediction has an error probability no larger than a user-specified risk level. To this end, we propose LEC, a principled framework that reframes selective prediction as a decision problem governed by a linear expectation constraint over selection and error indicators. This formulation directly controls the ratio between the expected number of accepted errors and the expected number of accepted predictions, which corresponds to the marginal error probability conditioned on selection. Under exchangeability, we derive a finite-sample sufficient condition that relies only on a held-out calibration set, enabling the computation of a risk-constrained, retention-maximizing threshold. Furthermore, we extend LEC to two-model routing systems: if the primary model's uncertainty exceeds its calibrated threshold, the input is delegated to a subsequent model, while maintaining system-level selection-conditioned error control. Experiments on both closed-ended and open-ended question answering (QA) and vision question answering (VQA) demonstrate that LEC maintains the prescribed risk level in accepted predictions and substantially improves sample retention compared to baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。