提出新评估方法,避免模型选择偏差导致的虚假优势
Selection-Aware Stress Testing for Interactive Agents
- 用发现任务特征学习任务权重,再在独立验证任务上测试
- 480轮实验中3.75分优势在验证时消失,证明原结论不可靠
- 适合需可靠评估的AI交互系统研发者,防止误判模型表现
当前智能体评估常采用单一基准选择工作流,再搜索其优势减弱的任务类型,但两者均来自相同数据,易产生选择偏差。本文提出选择感知语义压力测试(SASST),从发现任务的预执行特征中学习任务重加权,并在独立的确认任务上评估相同的配对比较。该协议检验支持度与稳定性,使用联合置信界覆盖所有计划声明,可返回无结论。在满足聚类假设下,证明了条件渐近有效性。40个聚类的审计显示高斯覆盖不足和保守的Bonferroni t边界。在一项包含480个回合的τ-基准研究中,3.75分的发现优势在确认阶段消失;另一模型研究也未确认任何工作流优势或稳定的压力规则。
原文摘要 · Abstract (English)
Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advantage weakens, so both conclusions are selected from the same data. We introduce Selection-Aware Semantic Stress Testing (\SASST{}), which learns a task reweighting from pre-execution features on discovery tasks and evaluates the same paired comparison on separate confirmation tasks. The protocol checks support and stability, uses joint bounds for all planned claims, and can return no claim. We prove conditional asymptotic validity under stated cluster assumptions. A forty-cluster audit finds Gaussian undercoverage and conservative Bonferroni $t$ bounds. In one 480-episode $τ$-bench study, a $3.75$ point discovery gain vanished on confirmation. A second-model study likewise confirmed neither a workflow benefit nor a stable stress rule.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。