arXiv:2504.20651stat.MLcs.LG2025-04被引 2

研究混合数据下的泛化能力,揭示了何时可将异质数据当作同质处理。

Learning and Generalization with Mixture Data

  • 用配对总变差刻画子群体差异,分析混合数据的学习可行性。
  • 在多种回归任务中获得泛化误差与收敛率的理论边界。
  • 发现模型越复杂,对数据异质性的容忍度越低,结论直观可信。

在众多机器学习应用中,训练数据天然存在异质性(如联邦学习、对抗攻击和神经网络中的域适应)。数据异质性被视为现代大规模学习的主要挑战之一。经典方法是通过混合模型表示异质数据。本文研究从混合分布采样的数据的泛化性能与统计速率。首先,我们以子群体分布之间的成对总变差距离刻画混合数据的异质性。进而,本文的核心主题是界定混合数据可被视作单一同质分布进行学习的范围。具体地,在经典 PAC 框架下研究泛化性能,并针对参数型(线性回归、超平面混合)及非参数型(Lipschitz、凸函数、Hölder 光滑)回归问题分析统计误差率。为此,我们推导了混合数据下的 Rademacher 复杂度与局部 Gaussian 复杂度界,并分别应用于获得泛化与收敛速率。结果表明,随着函数类复杂度提升,对成对总变差距离的要求也更严格,符合直觉。此外,对混合线性回归情形进行了精细分析,给出了基于异质性的紧致泛化误差上界。

原文摘要 · Abstract (English)

In many, if not most, machine learning applications the training data is naturally heterogeneous (e.g. federated learning, adversarial attacks and domain adaptation in neural net training). Data heterogeneity is identified as one of the major challenges in modern day large-scale learning. A classical way to represent heterogeneous data is via a mixture model. In this paper, we study generalization performance and statistical rates when data is sampled from a mixture distribution. We first characterize the heterogeneity of the mixture in terms of the pairwise total variation distance of the sub-population distributions. Thereafter, as a central theme of this paper, we characterize the range where the mixture may be treated as a single (homogeneous) distribution for learning. In particular, we study the generalization performance under the classical PAC framework and the statistical error rates for parametric (linear regression, mixture of hyperplanes) as well as non-parametric (Lipschitz, convex and Hölder-smooth) regression problems. In order to do this, we obtain Rademacher complexity and (local) Gaussian complexity bounds with mixture data, and apply them to get the generalization and convergence rates respectively. We observe that as the (regression) function classes get more complex, the requirement on the pairwise total variation distance gets stringent, which matches our intuition. We also do a finer analysis for the case of mixed linear regression and provide a tight bound on the generalization error in terms of heterogeneity.

泛化理论混合模型异质数据统计学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。