用二元监督训练模型无法同时获得准确且多样化的置信度估计。
Disproving the Feasibility of Learned Confidence Calibration Under Binary Supervision: An Information-Theoretic Impossibility
- 二元正确/错误标签无法区分不同置信度的正确预测。
- 实验显示负奖励导致置信度极端偏低(ECE>0.8),多样性不足(std<0.05)。
- 所有方法在真实数据集上失败率100%,说明这是理论限制而非技术问题。
我们证明了一个根本性不可能定理:神经网络在使用二元正确/错误监督训练时,无法同时学习到校准良好且具有意义多样性的置信度估计。通过严格的数学分析和涵盖负奖励训练、对称损失函数及事后校准方法的全面实证评估,我们揭示这是信息论上的约束,而非方法缺陷。实验显示,负奖励导致极端低估置信度(ECE > 0.8)并破坏置信度多样性(std < 0.05),对称损失无法摆脱二元信号平均化,而事后校准仅通过压缩置信度分布实现校准(ECE < 0.02)。我们将其形式化为一个欠定映射问题:正确预测的60%置信度与90%置信度收到相同监督信号。关键的是,真实世界验证表明,在MNIST、Fashion-MNIST和CIFAR-10上所有训练方法均100%失败,而事后校准33%的成功率反而印证了该定理——其成功依赖于变换而非学习。这一不可能性直接解释了神经网络幻觉现象,并确立事后校准在数学上是必要而非便利的。我们提出基于集成分歧和自适应多智能体学习的新监督范式,有望突破此根本限制,无需人工置信度标注。
原文摘要 · Abstract (English)
We prove a fundamental impossibility theorem: neural networks cannot simultaneously learn well-calibrated confidence estimates with meaningful diversity when trained using binary correct/incorrect supervision. Through rigorous mathematical analysis and comprehensive empirical evaluation spanning negative reward training, symmetric loss functions, and post-hoc calibration methods, we demonstrate this is an information-theoretic constraint, not a methodological failure. Our experiments reveal universal failure patterns: negative rewards produce extreme underconfidence (ECE greater than 0.8) while destroying confidence diversity (std less than 0.05), symmetric losses fail to escape binary signal averaging, and post-hoc methods achieve calibration (ECE less than 0.02) only by compressing the confidence distribution. We formalize this as an underspecified mapping problem where binary signals cannot distinguish between different confidence levels for correct predictions: a 60 percent confident correct answer receives identical supervision to a 90 percent confident one. Crucially, our real-world validation shows 100 percent failure rate for all training methods across MNIST, Fashion-MNIST, and CIFAR-10, while post-hoc calibration's 33 percent success rate paradoxically confirms our theorem by achieving calibration through transformation rather than learning. This impossibility directly explains neural network hallucinations and establishes why post-hoc calibration is mathematically necessary, not merely convenient. We propose novel supervision paradigms using ensemble disagreement and adaptive multi-agent learning that could overcome these fundamental limitations without requiring human confidence annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。