arXiv:2602.00286cs.LG2026-02被引 6

从信息论角度解析扩散模型生成顺序与并行解码的失败机制

Generation Order and Parallel Decoding in Masked Diffusion Models: An Information-Theoretic Perspective

  • 用信息论框架拆解生成顺序敏感性和并行化偏差
  • 发现易先解码能缓解模型误差,但并行采样会引入不可控错误
  • 验证了纠错需指数级代价,启发高效可靠的生成策略

掩码扩散模型(MDMs)通过牺牲序列确定性显著加速推理,但生成顺序机制与并行化风险的理论基础仍不清晰。本文提出统一的信息论框架,分离并分析两大失效根源:顺序敏感性与并行化偏差。核心发现:(1) 易先解码(优先处理低熵标记)在模型误差增大时优势更明显;(2) 因子化并行解码引入固有采样误差,可能导致逆KL散度任意增大,揭示标准前向KL忽略的“不连贯”故障;(3) 验证虽可消除采样误差,但代价呈指数增长,受块内总相关性约束;而重掩码等启发式方法虽高效,却无法保证分布正确性。在控制性块-马尔可夫模型和大规模MDM(LLaDA)算术推理任务上的实验验证了理论框架的有效性。

原文摘要 · Abstract (English)

Masked Diffusion Models (MDMs) significantly accelerate inference by trading off sequential determinism. However, the theoretical mechanisms governing generation order and the risks inherent in parallelization remain under-explored. In this work, we provide a unified information-theoretic framework to decouple and analyze two fundamental sources of failure: order sensitivity and parallelization bias. Our analysis yields three key insights: (1) The benefits of Easy-First decoding (prioritizing low-entropy tokens) are magnified as model error increases; (2) factorized parallel decoding introduces intrinsic sampling errors that can lead to arbitrary large Reverse KL divergence, capturing "incoherence" failures that standard Forward KL metrics overlook; and (3) while verification can eliminate sampling error, it incurs an exponential cost governed by the total correlation within a block. Conversely, heuristics like remasking, though computationally efficient, cannot guarantee distributional correctness. Experiments on a controlled Block-HMM and large-scale MDMs (LLaDA) for arithmetic reasoning validate our theoretical framework.

扩散模型生成顺序并行解码信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。