K路能量探针在判别式预测编码网络中实际等价于Softmax,无法提供更优信号。
K-Way Energy Probes for Metacognition Reduce to Softmax in Discriminative Predictive Coding Networks

- 通过近似推导证明能量探针可分解为Log-Softmax加一个无关正确性的残差项
- 六种实验条件下探针性能始终低于Softmax,且差距稳定在10^-3以内
- 适用于研究预测编码网络内部机制的学者,尤其关注结构探针有效性者
我们提出一个负结果及其解释机制,而非正式上界。预测编码网络(PCNs)支持一种K路能量探针:固定每个候选类别为目标,运行推理至收敛,比较各假设的最终能量。该探针看似比Softmax读取更丰富信号,因每假设能量依赖完整生成链。但在标准Pinchetti风格判别式PC框架下,此现象具有误导性。我们给出近似简化,表明在目标固定、交叉熵能量训练及有效前馈潜空间动力学条件下,K路能量边界可分解为单调函数的Log-Softmax边界加一个未训练相关于正确性的残差项。该分解预测结构探针应从下方跟踪Softmax。我们在CIFAR-10上测试了六个条件:扩展确定性训练、推理时潜变量运动直接测量、反向解码器公平性控制、匹配预算的PC与BP对比、五点朗之万温度扫描、轨迹积分MCPC训练。所有条件下探针均低于Softmax,差距在判别式PC家族内训练方式间保持稳定。最终状态与轨迹积分训练所得探针在确定性评估下AUROC_2差异小于10^-3。实验规模较小:单随机种子,210万参数网络,1280张测试图像。将结果作为预印本,邀请复现。讨论了分解不适用情形(双向PC、前瞻性配置、生成式PC、非CE能量形式),并指出分析未排除的有前景结构探针方向。
原文摘要 · Abstract (English)
We present this as a negative result with an explanatory mechanism, not as a formal upper bound. Predictive coding networks (PCNs) admit a K-way energy probe in which each candidate class is fixed as a target, inference is run to settling, and the per-hypothesis settled energies are compared. The probe appears to read a richer signal source than softmax, since the per-hypothesis energy depends on the entire generative chain. We argue this appearance is misleading under the standard Pinchetti-style discriminative PC formulation. We present an approximate reduction showing that with target-clamped CE-energy training and effectively-feedforward latent dynamics, the K-way energy margin decomposes into a monotone function of the log-softmax margin plus a residual that is not trained to correlate with correctness. The decomposition predicts that the structural probe should track softmax from below. We test this across six conditions on CIFAR-10: extended deterministic training, direct measurement of latent movement during inference, a post-hoc decoder fairness control on a backpropagation network, a matched-budget PC vs BP comparison, a five-point Langevin temperature sweep, and trajectory-integrated MCPC training. In every condition the probe sat below softmax. The gap was stable across training procedures within the discriminative PC family. Final-state and trajectory-integrated training produced probes whose AUROC_2 values differed by less than 10^-3 at deterministic evaluation. The empirical regime is small: single seed, 2.1M-parameter network, 1280 test images. We frame the result as a preprint inviting replication. We discuss conditions under which the decomposition does not apply (bidirectional PC, prospective configuration, generative PC, non-CE energy formulations) and directions for productive structural probing the analysis does not foreclose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。