给KAN网络加曲率惩罚,让模型更平滑可解释
KANs need curvature: penalties for compositional smoothness
- 提出无基底依赖的曲率惩罚机制
- 使模型保持精度同时激活函数更平滑
- 适合需要可解释性的科学机器学习场景
Kolmogorov-Arnold网络(KANs)凭借可学习的一元激活函数组合,在准确性和可解释性之间实现了强大平衡。然而,拟合良好的KAN模型其激活函数常表现出病态高曲率振荡,难以解释,且标准正则化无法抑制此现象。本文推导出一种无基底依赖的曲率惩罚,并证明施加该惩罚后模型可在保持高精度的同时实现显著更平滑的激活函数。通过分析函数组合如何影响曲率,我们给出了整个模型曲率相对于曲率惩罚的上界,并据此提出更丰富的惩罚形式。随着科学机器学习日益受限于精度与可解释性之间的权衡,此类在不牺牲精度的前提下提升可解释性的成果,将进一步增强KAN作为预测与洞察工具的实际应用价值。
原文摘要 · Abstract (English)
Kolmogorov-Arnold networks (KANs) offer a potent combination of accuracy and interpretability, thanks to their compositions of learnable univariate activation functions. However, the activations of well-fitting KANs tend to exhibit pathologically high-curvature oscillations, making them difficult to interpret, and standard regularization penalties do not prevent this. Here we derive a basis-agnostic curvature penalty and show that penalized models can maintain accuracy while achieving substantially smoother activations. Accounting for how function composition shapes curvature, we prove an upper bound on the full model's curvature relative to the curvature penalty, and use this to motivate richer forms of penalties. Scientific machine learning is increasingly bottlenecked by the trade-off between accuracy and interpretability. Results such as ours that improve interpretability without sacrificing accuracy will further strengthen KANs as a practical tool for both prediction and insight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。