arXiv:2605.28001cs.AIcs.CR2026-05

检验锚定解码中k-NAF预算机制的有效性,发现实际消耗远低于设定预算。

An Empirical Audit of k-NAF Budget Accounting for Anchored Decoding

  • 通过固定与自适应两种测试方式评估预算机制
  • 平均累积KL消耗远低于设定的600或1000预算值
  • 高预算比例现象多为代理指标偏差,非真实预算耗尽

我们通过两种方法实证检验锚定解码中的k-NAF预算会计机制:(i)固定类分层工作负载(约8500次随机执行,覆盖六类提示);(ii)针对高代理支出比率的自适应提示搜索。在固定负载下,平均累积KL支出远低于序列级预算K ∈ {600, 1000},且经验伯恩斯坦型代理支出始终低于K,表面重叠诊断(ROUGE-L与5-gram Jaccard)也较小。自适应搜索虽提升代理支出比率,但未导致明显预算耗尽。在持有集版权域数据上,k=3时部分提示在早期停止评估下代理比高于1,但扩大采样量后代理比降至[0.26, 0.40]区间,平均支出相近,表明该现象更可能源于代理指标偏差而非轨迹级预算失败。

原文摘要 · Abstract (English)

We empirically audit the k-NAF budget-accounting mechanism in Anchored Decoding using (i) a fixed, class-stratified workload (approximately 8,500 randomized executions across six prompt classes) and (ii) an adaptive prompt-search procedure targeting high proxy spend ratios. On the fixed workload, mean cumulative KL spend remains far below the sequence-level budgets K in {600, 1000}, and an empirical Bernstein-style proxy stays below K for every class; surface-overlap diagnostics (ROUGE-L and 5-gram Jaccard) are correspondingly small. Adaptive search increases the proxy spend ratio but does not produce clear budget exhaustion. On a held-out copyright-domain workload at k = 3, several prompts exhibit proxy ratios above 1 under early-stopped evaluations with small realized sample sizes; re-evaluating the same prompts with larger allocation reduces the proxy ratio to the range [0.26, 0.40] under comparable mean spend, consistent with proxy artifacts rather than per-trajectory budget failures.

解码算法预算控制实验分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。