arXiv:2605.19779cs.AIcs.LG2026-05中稿 · ICML

为连续智能体评估提供无需分布假设的不确定性量化方法

Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation

论文配图:Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation
图 1 · 摘自论文原文
  • 采用分割置信预测与自适应置信推断,实现无分布假设的覆盖保证
  • 24小时预测误差校准低于0.02,新发布后区间自动扩宽35%并恢复收敛
  • 适用于多智能体流水线、排名对比及排行榜大规模测试场景

我们将分割置信预测和自适应置信推断(ACI)应用于连续智能体评估,为预测的质量评分提供无分布假设的覆盖保证。在24小时预测时域上,置信区间校准误差低于0.02,且在智能体发布后ACI能正确将区间扩宽35%,随后逐渐收敛。我们进一步构建了多智能体流水线的组合不确定性边界(在阶段间相关性ρ∈[-0.5, 0.9]的模拟中验证),设计了控制错误排名率的置信拒绝规则,并提出针对排行榜规模多重检验的FDR修正拒绝策略。基于每小时采集的18个实时信号对50个智能体进行评估,结果显示每个智能体的条件覆盖率集中在名义水平附近(均值80.4%,90%的智能体位于[72%, 90%]区间),跨源情感分歧可预测排名不稳定性(r=0.64, p<0.01)。环形控制验证表明该框架捕捉到了基准之外的信号(rho_s=0.52, p<0.01, n=35)。代码与数据已按CC BY 4.0开源。

原文摘要 · Abstract (English)

We adapt split conformal prediction and adaptive conformal inference (ACI) to continuous AI agent evaluation, providing distribution-free coverage guarantees for forecasted quality scores. Conformal intervals achieve calibration error below 0.02 across all nominal levels at the 24h horizon, while ACI correctly widens intervals by 35% following agent releases then reconverges. We further develop compositional uncertainty bounds for multi-agent pipelines (validated via simulation across inter-stage correlations rho in [-0.5, 0.9]), a conformal abstention rule for pairwise rankings with controlled false-ranking rate, and FDR-corrected abstention for leaderboard-scale multiple testing. Evaluating 50 agents via 18 real-time signals collected hourly, we show that per-agent conditional coverage is well-concentrated around the nominal level (mean 80.4%, 90% of agents within [72%, 90%]), and that cross-source sentiment divergence predicts ranking instability (r=0.64, p<0.01). A circularity-controlled validation confirms the framework captures signal beyond benchmarks (rho_s=0.52, p<0.01, n=35). Code and data are released under CC BY 4.0.

不确定性量化智能体评估置信推断多智能体系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。