arXiv:2506.16704cs.LGstat.ML2025-06NeurIPS被引 2

提出新维度衡量域泛化所需最少采样域数,理论严谨。

How Many Domains Suffice for Domain Generalization? A Tight Characterization via the Domain Shattering Dimension

  • 引入域破碎维数刻画域泛化样本复杂度
  • 证明该维度与经典VC维存在紧致数量关系
  • 为域泛化提供可计算的理论指导,适合研究者参考

我们研究了域泛化中的一个基础问题:给定一组域(即数据分布),需要从其中随机采样多少个域才能收集足够数据,以训练出在所有已见和未见域上表现良好的模型?本文在PAC框架下建模此问题,并提出一种新的组合度量——域破碎维数。我们证明该维度能精确刻画域样本复杂度。进一步地,建立了域破碎维数与经典VC维之间的紧致数量关系,表明标准PAC设定下可学习的任意假设类,在本文设定下同样可学习。

原文摘要 · Abstract (English)

We study a fundamental question of domain generalization: given a family of domains (i.e., data distributions), how many randomly sampled domains do we need to collect data from in order to learn a model that performs reasonably well on every seen and unseen domain in the family? We model this problem in the PAC framework and introduce a new combinatorial measure, which we call the domain shattering dimension. We show that this dimension characterizes the domain sample complexity. Furthermore, we establish a tight quantitative relationship between the domain shattering dimension and the classic VC dimension, demonstrating that every hypothesis class that is learnable in the standard PAC setting is also learnable in our setting.

域泛化理论分析机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。