用多项式表示法量化神经网络的简单性,提升泛化能力。
Quantifying and Optimizing Simplicity via Polynomial Representations

- 通过正交多项式基逼近网络在数据路径上的行为,得到低维函数表示。
- 多项式有效阶数能预测不同任务和架构的泛化性能,优于传统度量。
- 可微分的简洁性正则项显著改善图像、文本分类及强化学习的泛化。
深度网络常偏好‘简单’解,这种简约性偏差被认为对泛化至关重要,但缺乏通用的定量衡量方法。本文引入多项式表示作为神经函数的分布感知、低维代理:通过正交多项式基近似网络在数据依赖插值路径上的预测行为,获得紧凑的函数表示。我们证明该表示的有效阶数可作为实用的简约性度量,能跨任务与架构预测泛化性能,且一致优于现有泛化代理(如尖锐度)。此外,多项式表示自然导出可微分的简洁性正则项,在图像分类、文本分类、对比视觉-语言模型微调及强化学习中均持续提升泛化表现。
原文摘要 · Abstract (English)
Deep networks often exhibit a preference for "simple" solutions, and such a simplicity bias is widely believed to play a key role in generalization. Yet a broadly applicable, quantitative measure of simplicity remains elusive. We introduce polynomial representations as a distribution-aware, low-dimensional surrogate for neural functions: we approximate a network's predictive behavior along data-dependent interpolation paths using orthogonal polynomial bases, yielding a compact functional representation. We show that the effective degree of this representation serves as a practical simplicity metric that is predictive of generalization across tasks and architectures, and consistently outperforms existing generalization proxies such as sharpness. Finally, polynomial representations naturally yield a differentiable simplicity regularizer, which consistently improves generalization in image and text classification, fine-tuning contrastive vision-language models, and reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。