揭示块生成中信息缺失与模型缺陷的根源,提出新评估方法
Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting
- 用信息下界分离并行生成中的信息丢失与模型误差
- 实测显示最佳提案每轮接受率最多71%,且一词后下界下降86%~100%
- 当前模型仍有43%~92%的拒绝来自可改进的模型能力
块生成器在单次前向传播中预估多个标记,但在早期目标标记未生成时就可能被拒绝。其拒绝原因包含两部分:块内路径信息缺失与可观测信息建模不充分。现有接受长度无法区分二者。本文通过引入信息下界——在特定条件顺序下的最小期望拒绝率,将两者分离;高于下界的拒绝即为模型差距。基于四个领域、四个开源目标模型及一个前沿API目标模型的轨迹分析,得出三项发现:第一,在Qwen3-4B上,全并行信息下界在最终位置达到0.286,即使最优提案也仅能实现71%的每槽接受率;第二,仅需一个已生成标记,即可消除86%至100%的该下界,此局部性亦被独立互信息分析验证;第三,当前生成器仍远高于其下界:最终槽的模型差距占DFlash拒绝的43%–64%,占DSpark理想条件拒绝的85%–92%。这些结果清晰划分了短程条件作用与生成质量的贡献。
原文摘要 · Abstract (English)
Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. Accepted length cannot distinguish them. We separate the two with an information floor, the minimum expected rejection at a specified conditioning order; rejection above this floor is the model gap. Estimating both from target rollouts across four domains, four open-weight targets, and a frontier API target yields three findings. First, the all-parallel floor reaches $0.286$ at the final slot on Qwen3-4B, limiting even the best proposal to $71\%$ per-slot acceptance. Second, one realised token removes $86$--$100\%$ of this floor, a locality also recovered by an independent mutual-information analysis. Third, current drafters remain far above their floors: the final-slot model gap accounts for $43$--$64\%$ of DFlash rejection and $85$--$92\%$ of DSpark's oracle-conditioned rejection. These findings separate the value of short-range conditioning from proposal quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。