arXiv:2602.00315cs.LGcs.AI2026-02

用精确后验揭示神经网络真实极限,发现训练仍在进步却常被误判。

Beyond the Loss Curve: Scaling Laws, Active Learning, and the Limits of Learning from Exact Posteriors

  • 用条件归一化流构建精确后验,突破传统评估局限
  • 数据量增大时错误持续下降,即使损失曲线已平缓
  • 可识别真正有价值样本,提升主动学习效率

神经网络离最优性能还有多远?标准评估因无法获取真实后验 p(y|x) 而难以回答。本文利用类别条件归一化流作为预言者,在真实图像数据集(AFHQ、ImageNet)上实现精确后验的可计算性,开启五项研究:1)缩放律分析显示预测误差由不可约的随机不确定性与可减少的认知误差构成,后者随数据量呈幂律下降,即便总损失已趋于平稳;2)可精确测量随机不确定性下限,不同架构表现差异显著——ResNet 呈现清晰幂律,而视觉变换器在小数据下出现停滞;3)使用精确后验训练优于硬标签,实现近乎完美的校准;4)预言者可计算受控扰动下的精确KL散度,表明分布偏移类型比幅度影响更大:类别失衡影响微弱,而输入噪声在特定散度值下引发灾难性退化;5)精确认知不确定性能区分真正信息量高的样本与固有模糊样本,显著提升样本效率。本框架揭示标准指标掩盖了持续学习、隐藏架构差异,并无法诊断分布偏移的本质。

原文摘要 · Abstract (English)

How close are neural networks to the best they could possibly do? Standard benchmarks cannot answer this because they lack access to the true posterior p(y|x). We use class-conditional normalizing flows as oracles that make exact posteriors tractable on realistic images (AFHQ, ImageNet). This enables five lines of investigation. Scaling laws: Prediction error decomposes into irreducible aleatoric uncertainty and reducible epistemic error; the epistemic component follows a power law in dataset size, continuing to shrink even when total loss plateaus. Limits of learning: The aleatoric floor is exactly measurable, and architectures differ markedly in how they approach it: ResNets exhibit clean power-law scaling while Vision Transformers stall in low-data regimes. Soft labels: Oracle posteriors contain learnable structure beyond class labels: training with exact posteriors outperforms hard labels and yields near-perfect calibration. Distribution shift: The oracle computes exact KL divergence of controlled perturbations, revealing that shift type matters more than shift magnitude: class imbalance barely affects accuracy at divergence values where input noise causes catastrophic degradation. Active learning: Exact epistemic uncertainty distinguishes genuinely informative samples from inherently ambiguous ones, improving sample efficiency. Our framework reveals that standard metrics hide ongoing learning, mask architectural differences, and cannot diagnose the nature of distribution shift.

神经网络极限主动学习后验估计缩放律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。