arXiv:2608.28011cs.AI2026-08

零阶扰动预算分配不提升大模型代理的样本效率

Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM Agents

  • 用失败归因机制分配扰动预算,但效果不如均匀分配
  • 在多个模型和任务上均未发现显著性能提升(<2%)
  • 揭示了三类隐蔽实验失效模式,适用于研究者警惕

轨迹级归因可仅通过可验证信号定位工具使用型大模型代理中导致失败的模块。我们探究此类失败归因是否应指导固定的零阶/进化策略(ZO/ES)扰动预算分配。在合成环境及冻结的Qwen2.5-1.5B/3B、SmolLM2-1.7B代理上,覆盖三个任务族、六种分配方案、信用噪声测试、成对种子与精确符号翻转实验,结果表明:在任意池内比较中,均无统计可检测到的改进(无至少2个百分点的增益)。联合软加总-σ方案在1.5B和3B模型上与均匀分配等效(AUC差异±0.02内);将全部预算集中于归因最大值模块在1.5B上略等效(该模块为已验证瓶颈),但在3B上显著更差。逆倾向去偏无法挽救路由效果,错误路由导致最高-0.074 AUC内部损失与-0.118端到端损失(来自BFCL衍生任务族)。在六种固定步数调度下,损失与瓶颈饥饿率呈线性关系(R²=0.94,描述性强),预注册的无信用覆盖率底限可消除检测到的损害。匹配预算爆发与步数补偿追赶调度一致表明,损害源于累积参数移动不足,而非更新频率。核心评估指标为固定任务池上的优化效率。在未见的BFCL函数上,唯一例外是软路由在保留端点上优于均匀分配(+0.047,p=0.031,n=6)。一种可能但未经验证的解释是:路由偏好的调用者改进得以传递,而均匀分配的池内增益反映的是特定于实验框架的合成器行为。本研究明确报告此例外,并记录三类可无声破坏ZO/ES实验的失效模式。

原文摘要 · Abstract (English)

Trajectory-level credit assignment can localize which module of a tool-using LLM agent causes failures using only verifiable signals. We ask whether such failure credit should route a fixed zeroth-order/evolution-strategies (ZO/ES) perturbation budget. Across a synthetic environment and frozen Qwen2.5-1.5B/3B and SmolLM2-1.7B agents, three task families, six allocation schemes, a credit-noise sweep, paired seeds, and exact sign-flip tests, we find no statistically detectable improvement over uniform allocation in any on-pool comparison (no gain of at least 2 percentage points). The joint soft-plus-sigma scheme is equivalent to uniform within a +/- 0.02 AUC margin on 1.5B and 3B; concentrating the full budget on the credit argmax is marginally equivalent on 1.5B, where that module is the verified bottleneck, and significantly worse on 3B. Inverse-propensity debiasing does not rescue routing, and misrouting costs up to -0.074 AUC in-house and -0.118 end-to-end on the BFCL-derived family. Across six fixed-step schedules, loss is linear in bottleneck starvation rate (R^2 = 0.94, descriptive), and a preregistered credit-free coverage floor removes detected harm. Matched-budget burst and step-compensating catch-up schedules are consistent with harm arising from insufficient cumulative parameter movement rather than update frequency. Our primary estimand is optimization efficiency on a fixed task pool. On unseen BFCL functions, the study's one exception is that soft routing exceeds uniform on held-out endpoints (+0.047, p = 0.031, n = 6). A plausible but untested reading is that routing-favored caller improvements transfer while uniform's on-pool gains reflect a synthesizer behavior specific to our harness. We report this exception explicitly and document three failure modes that can silently invalidate ZO/ES experiments on frozen LLMs.

大模型优化效率实验设计零阶优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。