arXiv:2608.22331cs.CL2026-08

评估大模型工具调用的噪声底限,发现提示扰动影响远大于重试差异。

Noise Floor Audit for Agent Benchmarks

  • 用一致语法树评分法对比三家厂商工具调用稳定性
  • 同一提示下重试结果几乎确定,但提示微调导致误差扩大11至58倍
  • 输出格式错误占失败30%,说明准确率掩盖了可靠性问题

我们对两个提供商的三个原生工具调用端点,在官方BFCL多任务与并行类别上,采用匹配的抽象语法树(AST)评分法进行了测量变异性审计。在温度0设置下,Groq端点和启用了思考功能的Gemini设置中,重试几乎确定:波动翻转比例分别为0.7%、2.0%和2.7%,平均运行相关性分别为0.997、0.966和0.961。语义保持的提示扰动在所有端点上造成更大的噪声底限,其配对标准差中位数是重试配对标准差的11到58倍。失败特征也发生变化:格式错误导致的失败分别占任务失败的30%、7%和不到1%,表明边际准确率不仅掩盖了稳定性问题,还隐藏了失败模式差异。

原文摘要 · Abstract (English)

We audit measurement variability for 3 native tool-calling endpoints across 2 providers on the official BFCL multiple and parallel categories, using matched AST grading. At temperature 0, reruns are nearly deterministic across Groq endpoints and a thinking-enabled Gemini setting: ever-flip fractions are 0.7%, 2.0%, and 2.7%, with mean run correlations of 0.997, 0.966, and 0.961. Semantics-preserving prompt perturbations create the larger floor on all endpoints, with median perturbation paired SDs 11x to 58x larger than rerun paired SDs. The failure character also shifts: malformed-output failures account for 30%, 7%, and <1% of task failures, so marginal accuracy hides not only stability but also failure mode.

模型评测工具调用稳定性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。