arXiv:2410.07427cs.LGstat.ML2024-10被引 6

提出隐式网络的泛化界,揭示其理论可靠性。

A Generalization Bound for a Family of Implicit Networks

  • 基于压缩映射的不动点定义网络结构
  • 通过覆盖数推导出泛化误差上界
  • 为隐式网络提供理论支撑,适合研究者参考

隐式网络是一类输出由参数化算子的不动点定义的神经网络,在自然语言处理、图像处理等多个领域取得成功。尽管其在实践中表现优异,但关于其泛化能力的理论研究仍不充分。本文研究一类由参数化压缩映射定义的隐式网络家族,基于Rademacher复杂度的覆盖数论证,给出了该类网络的泛化误差上界,为隐式网络提供了理论支持。

原文摘要 · Abstract (English)

Implicit networks are a class of neural networks whose outputs are defined by the fixed point of a parameterized operator. They have enjoyed success in many applications including natural language processing, image processing, and numerous other applications. While they have found abundant empirical success, theoretical work on its generalization is still under-explored. In this work, we consider a large family of implicit networks defined parameterized contractive fixed point operators. We show a generalization bound for this class based on a covering number argument for the Rademacher complexity of these architectures.

隐式网络泛化界理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。