arXiv:2505.22622stat.MLcs.LG2025-05

发现简单性是模型泛化能力的关键,提出理论框架提升模型在分布外场景的表现。

Principled Out-of-Distribution Generalization via Simplicity

  • 基于简单性度量,构建分布外泛化的理论框架。
  • 在固定差距和渐近差距两种设定下,给出首个精确的样本复杂度保证。
  • 适合关注模型泛化机制与理论分析的研究者。

现代基础模型展现出卓越的分布外(OOD)泛化能力,能解决远超训练数据支持范围的任务。然而,其背后的理论原理仍不清晰。本文通过分析扩散模型在图像生成中的组合泛化能力,发现尽管神经网络架构足够表达多种模型(包括对分布外输入表现不佳的模型),但符合人类预期且具有可泛化能力的真实模型,通常是在与训练数据一致的所有模型中最为简单的那个。受此启发,我们提出了基于简单性的分布外泛化理论框架,采用预定义的简单性度量进行量化。分析了两种关键情形:(1) 固定差距情形,即真实模型比所有伪相关模型严格更简单,且差距恒定;(2) 渐近差距情形,用平滑性条件替代固定差距,确保与真实模型简单性相近的模型预测结果相似。针对这两种情形,研究了正则化最大似然估计器,并首次建立了学习真实、可泛化、简单模型的尖锐样本复杂度保证。

原文摘要 · Abstract (English)

Modern foundation models exhibit remarkable out-of-distribution (OOD) generalization, solving tasks far beyond the support of their training data. However, the theoretical principles underpinning this phenomenon remain elusive. This paper investigates this problem by examining the compositional generalization abilities of diffusion models in image generation. Our analysis reveals that while neural network architectures are expressive enough to represent a wide range of models -- including many with undesirable behavior on OOD inputs -- the true, generalizable model that aligns with human expectations typically corresponds to the simplest among those consistent with the training data. Motivated by this observation, we develop a theoretical framework for OOD generalization via simplicity, quantified using a predefined simplicity metric. We analyze two key regimes: (1) the constant-gap setting, where the true model is strictly simpler than all spurious alternatives by a fixed gap, and (2) the vanishing-gap setting, where the fixed gap is replaced by a smoothness condition ensuring that models close in simplicity to the true model yield similar predictions. For both regimes, we study the regularized maximum likelihood estimator and establish the first sharp sample complexity guarantees for learning the true, generalizable, simple model.

泛化能力扩散模型理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。