提出新诊断方法,检测VAE中编码与解码的匹配度。
Lost and Found in Translation: Variational Diagnostics for Neural Codebook Channels

- 构建编码-解码通道矩阵,量化两者间信息传递质量。
- 在4个数据集上验证,95%以上情况满足理论边界约束。
- 适合关注模型可解释性与生成质量的深度学习研究者。
经典通信系统不仅因随机噪声失效,也因收发端使用不兼容的代码本而失败。变分自编码器(VAE)联合训练编码器 $q_ϕ$ 和解码器 $p_θ$,实践者常将潜在空间视为离散代码,用于聚类、条件生成和机制可解释性分析。然而,传统诊断指标——如ELBO、活跃单元数、互信息和代码直方图——仅能确认该代码是否被使用,无法判断解码器是否正确读取编码器的代码。本文提出神经代码本通道 $K_{e o d}(jackslash i)$,一种耦合编码器-解码器的诊断工具,其非对角线质量受一个无架构依赖的伯努利-KL 证书 $d_{\mathrm{bin}}(1-\mathcal{A} \| \barη_p) \le \barΔ$ 控制,该证书由变分间隙决定。证书是经典KL链式法则在编码-解码不一致事件上的操作化表达,并辅以构造性边际不可能性结果:任何边际直方图、熵、活跃代码数或互信息组合都无法确定 $K_{e\to d}$。我们在四个sklearn数据集(有限网格精确计算,5/5种子,20/20对满足边界)、二维模型(边界在观察到的不一致处2.71倍时仍非平凡,四重恒等式闭合误差小于 $10^{-4}$)、带重要性采样控制的MNIST以及达到预测极限 $\ ilde{\mathcal{A}}=1.000$ 的VQ-VAE上进行了审计。完整报告单元 $(K_{e\to d}, \mathcal{A}, R_{\mathrm{eff}}, R, \mathrm{AU})$ 可直接用于模型审计。更广泛地,该框架使经典通信理论中命名数十年的‘不匹配解码’故障模式,在单一深度生成模型内部变得可见。
原文摘要 · Abstract (English)
Classical communication systems fail not only through random noise but also when transmitter and receiver use incompatible operational codebooks. Variational autoencoders (VAEs) train an encoder $q_ϕ$ and decoder $p_θ$ jointly, and practitioners treat the resulting latent space as a discrete code -- for clustering, conditional generation, and mechanistic interpretability. Yet standard VAE diagnostics -- ELBO, active units, mutual information, and code histograms -- certify only whether this code is used, never whether the decoder reads each latent under the encoder's code. We close this gap with the neural codebook channel $K_{e\to d}(j\mid i)$, a coupled encoder-decoder diagnostic whose off-diagonal mass is bounded by an architecture-free Bernoulli-KL certificate $d_{\mathrm{bin}}(1-\mathcal{A} \,\|\, \barη_p) \le \barΔ$ controlled by the variational gap. The certificate is the operational specialization of the classical KL chain rule under disintegration to the encoder-decoder disagreement event, complemented by a constructive marginal-impossibility result: no combination of marginal histograms, entropies, active-code counts, or mutual information determines $K_{e\to d}$. We audit the certificate on four sklearn datasets (finite-grid exact, 5/5 seeds, 20/20 pairs satisfy the bound), a 2D model where the bound is non-vacuous at $2.71\times$ the observed disagreement and the four-term identity closes within $10^{-4}$, MNIST under importance-sampling control, and a VQ-VAE attaining the predicted limit $\hat{\mathcal{A}}=1.000$. The package $(K_{e\to d}, \mathcal{A}, R_{\mathrm{eff}}, R, \mathrm{AU})$ is an audit-ready reporting unit. More broadly, the framework makes mismatched decoding -- a failure mode classical communication theory named decades ago -- visible inside a single deep generative model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。