arXiv:2606.30705cs.LGcs.AI2026-06

连续文本生成失败因解码器过锐,与训练无关。

Why Do Few-Step Text Latents Fail When Image Latents Work? Non-Commitment at Sharp Categorical Readouts

论文配图:Why Do Few-Step Text Latents Fail When Image Latents Work? Non-Commitment at Sharp Categorical Readouts
图 1 · 摘自论文原文
  • 解码器读出过锐导致离散分支无法分辨,引发生成崩溃。
  • 实测显示文本解码器读出锐度达10^5,图像解码器仅约1。
  • 自回归或随机重注可绕过此瓶颈,适合改进生成模型设计。

确定性少步生成在连续图像隐空间中表现良好,但在连续文本隐空间中会退化为无意义输出。我们证明其根本原因是几何限制而非训练或扩展缺陷:平滑且规则受限的确定性映射无法在尖锐分类读出前分辨离散分支选择。少步失败由解码器锐度决定,而非传输精度。在真实文本自编码器的重叠区域,我们证明(定理3)后验均值终点步骤的词元翻转率等于隐空间质量在决策边界附近$O(s(t))$管状区域内的分布率。两个诊断指标DABI(读出锐度)和CCI(分类承诺)在多个公开检查点上的测量显示,四种独立构建的连续文本解码器对边界对齐扰动的放大程度远超范数匹配的各向同性扰动(DABI从$5\times10^{2}$增至$>10^{5}$),而图像解码器的DABI≈1。两种机制可突破连续约束:分类承诺(自回归解码器即使在更锐读出下仍成功)和随机重注入(确定性ODE在K=4时PPL为294,而同一模型的SDE为50)。在理想分离区域,我们证明了匹配的尖锐传输规律,包括维度相图:将M个模式分离所需的确定性刚度在潜维数达到$Ω(\log M)$后增长为$Θ(\sqrt{\log M})$(固定维度下为$M^{1/n}$),深度为B的层次结构使每步峰值减小$\sqrt{B}$倍(定理5-7);共面积恒等式将这些结果与重叠管状区域关联(定理17)。结论是:在确定性连续类中存在不可回避的精度-深度-刚度权衡,两种逃逸方式均需脱离该类别。

原文摘要 · Abstract (English)

Deterministic few-step generation succeeds on continuous image latents but collapses to incoherent text on continuous text latents, and we show the cause is geometric rather than a training or scaling deficiency: a smooth, regularity-limited deterministic map cannot resolve a discrete branch choice before a sharp categorical readout, so few-step failure is governed by decoder sharpness, not transport accuracy. In the overlapping regime of real text autoencoders, we prove (Theorem 3) that the posterior-mean terminal step flips tokens at the rate of the latent mass in an $O(s(t))$ tube around decision boundaries. Two diagnostics, DABI (readout sharpness) and CCI (categorical commitment), measured on published checkpoints show that four independently built continuous-text decoders amplify a boundary-aligned perturbation far beyond a norm-matched isotropic one (DABI from $5\times10^{2}$ to $>10^{5}$), while image decoders have DABI $\approx 1$. Two mechanisms escape the continuous bound: categorical commitment (autoregressive decoders succeed despite sharper readouts) and stochastic re-injection (deterministic ODE at $K=4$ gives PPL 294 versus SDE 50 on the same model). In the idealized separated regime we prove matching sharp transport laws, including a dimension phase diagram: the deterministic stiffness needed to separate $M$ modes grows as $Θ(\sqrt{\log M})$ once the latent dimension is $Ω(\log M)$ (and as $M^{1/n}$ in fixed dimension), with a depth-$B$ hierarchy giving a $\sqrt{B}$-smaller per-step peak (Theorems 5-7); a coarea identity links these to the overlapping tube (Theorem 17). The result is an accuracy-depth-stiffness tradeoff: within the deterministic-continuous class the cost is irreducible, and both escapes step outside it.

生成模型隐空间解码器锐度文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。