对比两种采样方法在联合能量模型上的表现,发现改进版无噪声预测-校正法并未优于传统SGLD。
Comparing SGLD and a fixed-noise Predictor-Corrector adaptation in canonical Joint Energy-Based Models on CIFAR-10

- 用固定噪声的预测-校正法替代SGLD进行训练与生成测试
- 分类准确率92.88%,生成质量FID为44.46,略逊于基准值
- 未发现明显优势,理论保证不适用于当前设置,适合研究采样器鲁棒性者阅读
联合能量模型(JEM)将分类与生成统一于单一网络,并支持分布外检测。标准JEM训练依赖随机梯度朗之万动力学(SGLD);而理论上更优的预测-校正(PC)采样器此前未在标准模型上系统复现。本文在不使用归一化层的WideResNet-28-10上独立运行两次,测试固定噪声的PC改进版本——以确定性梯度步替代退火噪声预测器——在三种协议下的表现:全程替换SGLD训练(115-132轮)、冷启动生成(FID)和多分布外检测(AUROC)。分类测试准确率达92.88%(基准92.9%),缓冲区FID为44.46(基准38.40)。记录到两类失败模式:训练后期灾难性发散,具典型异常样本缓冲特征(四次运行均出现);以及针对SVHN的分布外识别动态受运行种子影响。各协议下均未观察到改进方法相对于SGLD的一致优势:精炼检测的AUROC差异低于0.007(十组检查点-分布外对);冷启动生成中SGLD领先约5个FID点;训练协议中,基于种子-图像分层抽样的宏观平均AUROC差异置信区间包含零,且双运行种子级等效检验无法确立形式等价性。训练数据既支持等价,也支持微小方向效应。结果符合理论预期:退火噪声下成立的PC保证无法转移至标准JEM的恒定噪声场景。
原文摘要 · Abstract (English)
Joint Energy-Based Models (JEM) unify classification and generation within a single network and support out-of-distribution (OOD) detection. Canonical JEM training relies on stochastic gradient Langevin dynamics (SGLD); a theoretically motivated alternative, the Predictor-Corrector (PC) sampler, has not previously undergone a systematic replication test on the canonical model. We reproduce canonical JEM on WideResNet-28-10 without normalisation layers on two independent runs and test a fixed-noise PC adaptation - with the degenerate annealed-noise predictor replaced by a deterministic gradient step - across three protocols: the adapted sampler replacing SGLD throughout the full training trajectories (115-132 epochs); cold-start generation (FID); and refinement-style multi-OOD detection (AUROC). The reconstruction reaches 92.88% test accuracy and buffer-FID 44.46 (canonical: 92.9% and 38.40). We document two failure modes: catastrophic late-training divergence with the signature of the canonical outlier-buffer mechanism (all four runs), and run-dependent SVHN OOD-discrimination dynamics. No consistent method-level advantage of the adaptation over SGLD is observed on any protocol: refinement AUROC differences stay below 0.007 across ten checkpoint-OOD pairs; seeded cold-start generation favours SGLD by about five FID points; on the training protocol a hierarchical seed-by-image bootstrap gives a 95% confidence interval on the macro-averaged AUROC difference that contains zero, while a seed-level equivalence test with two runs per method cannot establish formal equivalence. The training-protocol data are consistent both with equivalence and with a small directional effect. This outcome is consistent with theory: the guarantees of the annealed-noise PC framework do not transfer to the constant-noise regime of canonical JEM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。