研究发现:训练与测试分布不一致时,反而可能提升泛化性能。
Limitations of Using Identical Distributions for Training and Testing When Learning Boolean Functions
- 允许为不同训练分布选择最优学习算法,突破传统同分布假设。
- 在存在单向函数的条件下,不同分布训练可带来更好泛化效果。
- 仅当目标函数具特定规律性时,同分布仍是最优选择。
当训练数据与测试数据分布不一致时,泛化问题变得更为复杂,引发诸多疑问。以往研究表明,对某些固定学习方法,使用与测试分布不同的训练分布反而能提升泛化能力。然而,这些结果未考虑针对每种训练分布可选择最优学习算法的情况,因而无法判断收益是来自分布差异本身,还是学习器次优所致。本文在完全一般性下解决此问题:即当学习器可针对训练分布最优适配时,是否始终应使训练分布与测试分布相同?令人惊讶的是,在假设单向函数存在的前提下,答案是否定的——分布匹配并非总是最优。不过,若对目标函数施加某些正则性条件,则在均匀分布情形下,标准结论得以恢复。
原文摘要 · Abstract (English)
When the distributions of the training and test data do not coincide, the problem of understanding generalization becomes considerably more complex, prompting a variety of questions. Prior work has shown that, for some fixed learning methods, there are scenarios where training on a distribution different from the test distribution improves generalization. However, these results do not account for the possibility of choosing, for each training distribution, the optimal learning algorithm, leaving open whether the observed benefits stem from the mismatch itself or from suboptimality of the learner. In this work, we address this question in full generality. That is, we study whether it is always optimal for the training distribution to be identical to the test distribution when the learner is allowed to be optimally adapted to the training distribution. Surprisingly, assuming the existence of one-way functions, we find that the answer is no. That is, matching distributions is not always the best scenario. Nonetheless, we also show that when certain regularities are imposed on the target functions, the standard conclusion is recovered in the case of the uniform distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。