给黑箱AI输出打可靠性分数,确保可信部署。
Black-Box Reliability Certification for AI Agents via Self-Consistency Sampling and Conformal Calibration
- 用自一致性采样和置信校准计算可靠性分数。
- GPT-4.1在GSM8K上达94.6%可靠性,更弱模型得分更低。
- 适合关注AI输出可信度的开发者与评估者。
针对黑箱AI系统与任务,我们提出一种可靠性水平——基于自一致性采样与置信校准,为每个系统-任务对生成一个单一数值,作为部署门槛。该方法提供精确、有限样本、无需分布假设的保证。自一致性采样可指数级降低不确定性;置信校准确保正确性覆盖范围在目标水平的1/(n+1)内,且不受系统误差影响。通过更大答案集揭示难度,使难题表现更透明。较弱模型获得更低可靠性(非准确率):GPT-4.1在GSM8K上为94.6%,TruthfulQA上为96.8%;而GPT-4.1-nano在GSM8K上为89.8%,在MMLU上为66.5%。在五个基准、三类模型家族及合成与真实数据上验证有效。可解项条件覆盖率达0.93以上;序列停止机制降低约50%的API成本。
原文摘要 · Abstract (English)
Given a black-box AI system and a task, at what confidence level can a practitioner trust the system's output? We answer with a reliability level -- a single number per system-task pair, derived from self-consistency sampling and conformal calibration, that serves as a black-box deployment gate with exact, finite-sample, distribution-free guarantees. Self-consistency sampling reduces uncertainty exponentially; conformal calibration guarantees correctness within 1/(n+1) of the target level, regardless of the system's errors -- made transparently visible through larger answer sets for harder questions. Weaker models earn lower reliability levels (not accuracy -- see Definition 2.4): GPT-4.1 earns 94.6% on GSM8K and 96.8% on TruthfulQA, while GPT-4.1-nano earns 89.8% on GSM8K and 66.5% on MMLU. We validate across five benchmarks, five models from three families, and both synthetic and real data. Conditional coverage on solvable items exceeds 0.93 across all configurations; sequential stopping reduces API costs by around 50%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。