arXiv:2511.09996cs.LGstat.ML2025-11

新学习范式突破大规模模型泛化瓶颈,无需预设参数即可利用数据内在结构。

A Novel Data-Dependent Learning Paradigm for Large Hypothesis Classes

  • 基于数据驱动构建学习框架,减少对先验假设的依赖
  • 在相似性、聚类、光滑性等常见假设下实现良好泛化性能
  • 适合缺乏先验知识的大规模模型学习场景

针对候选模型集合过大导致经验损失难以统一收敛的问题,传统方法多采用结构风险最小化(SRM)或正则化策略。本文提出一种新型数据依赖学习范式,更充分地融合实证数据,降低算法决策对先验假设的依赖。我们分析了该方法的泛化能力,并在多种典型学习假设下验证其有效性,包括近邻点相似性、域内聚类形成高度标签同质区域、标签函数的Lipschitz连续性以及对比学习假设。该方法可在不预先知晓这些假设真实参数的情况下加以利用,显著提升复杂模型空间的学习效率与稳定性。

原文摘要 · Abstract (English)

We address the general task of learning with a set of candidate models that is too large to have a uniform convergence of empirical estimates to true losses. While the common approach to such challenges is SRM (or regularization) based learning algorithms, we propose a novel learning paradigm that relies on stronger incorporation of empirical data and requires less algorithmic decisions to be based on prior assumptions. We analyze the generalization capabilities of our approach and demonstrate its merits in several common learning assumptions, including similarity of close points, clustering of the domain into highly label-homogeneous regions, Lipschitzness assumptions of the labeling rule, and contrastive learning assumptions. Our approach allows utilizing such assumptions without the need to know their true parameters a priori.

学习范式泛化能力数据依赖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。