揭示回归中插值与聚合的相互作用,找到最优学习方法
The Interplay Between Interpolation and Aggregation in Regression: Optimal Sample Complexity
- 用γ-图维数刻画聚合学习的可学习性
- 仅用三个插值假设取中位数即达最优性能
- 某些问题需无限聚合或非插值规则才能学习
本文从理论上研究了回归中插值与聚合的相互作用。我们证明γ-图维数能刻画一大类自然聚合过程的可学习性。进一步证明,一种极简的聚合方法——通过中位数组合三个插值假设,是所有此类聚合方法中的最优解,且严格优于正则学习。最后表明,某些假设类仅能通过聚合无限多个假设或使用非插值聚合规则(可能预测超出输入范围)来学习,任何有限插值聚合均无法达到基本性能。
原文摘要 · Abstract (English)
This work investigates theoretically the interplay between interpolation and aggregation in regression. We establish that the $γ$-graph dimension characterizes learnability for a broad class of natural aggregation procedures. Furthermore, we prove that an extremely simple aggregation procedure, combining three interpolating hypotheses via the median, is optimal among all these aggregation procedures, and is strictly more powerful than proper learning. Finally, we show that some hypothesis classes are learnable only by aggregating infinitely many hypotheses or by using non-interpolating aggregation rules (which may predict outside the range of their inputs), and any finite interpolating aggregation fails to achieve even trivial performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。