提出温度损失提升KAN网络概率预测可靠性
PostHoc FREE Calibrating on Kolmogorov Arnold Networks
- 引入可学习温度参数的损失函数,动态调节预测分布
- 实测显示该方法显著降低校准误差,提升置信度准确性
- 适合关注模型可信度的从业者,尤其在数据稀疏区有效
Kolmogorov Arnold Networks(KANs)基于柯尔莫哥洛夫-阿诺德表示定理,采用B样条参数化实现灵活的局部自适应函数逼近。尽管能捕捉超越标准多层感知机(MLPs)的复杂非线性关系,但其经常出现置信度校准偏差——在密集数据区域过度自信,在稀疏区域则低估。本文系统分析了层宽、网格阶数、捷径函数和网格范围四个关键超参数对校准性能的影响,并提出一种新型温度缩放损失(TSL),将温度参数直接嵌入训练目标,动态调整学习过程中的预测分布。理论分析与在标准基准上的广泛实验表明,TSL显著降低校准误差,从而提高概率预测的可靠性。本研究为样条类神经网络的设计提供了可操作洞见,并确立了TSL作为增强校准能力的稳健损失方案。
原文摘要 · Abstract (English)
Kolmogorov Arnold Networks (KANs) are neural architectures inspired by the Kolmogorov Arnold representation theorem that leverage B Spline parameterizations for flexible, locally adaptive function approximation. Although KANs can capture complex nonlinearities beyond those modeled by standard MultiLayer Perceptrons (MLPs), they frequently exhibit miscalibrated confidence estimates manifesting as overconfidence in dense data regions and underconfidence in sparse areas. In this work, we systematically examine the impact of four critical hyperparameters including Layer Width, Grid Order, Shortcut Function, and Grid Range on the calibration of KANs. Furthermore, we introduce a novel TemperatureScaled Loss (TSL) that integrates a temperature parameter directly into the training objective, dynamically adjusting the predictive distribution during learning. Both theoretical analysis and extensive empirical evaluations on standard benchmarks demonstrate that TSL significantly reduces calibration errors, thereby improving the reliability of probabilistic predictions. Overall, our study provides actionable insights into the design of spline based neural networks and establishes TSL as a robust loss solution for enhancing calibration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。