arXiv:2501.03858cs.LGstat.ML2025-01被引 8

对称性提升模型泛化能力,证明了不变性和等变性是有效归纳偏置。

Symmetry and Generalisation in Machine Learning

  • 用平均算子证明非等变预测器总有更低风险的等变版本。
  • 在随机设计最小二乘和核岭回归中量化了对称性带来的泛化性能提升。
  • 适合研究模型归纳偏置与对称性关系的学者参考。

本文研究不变性和等变性在监督学习中对泛化的影响。通过引入平均算子视角,我们证明:对于任何不满足等变性的预测器,总存在一个等变预测器,在所有等变性正确设定的回归问题上具有严格更低的测试风险。这为对称性作为归纳偏置的有效性提供了严谨证明。我们将该思想应用于随机设计下的最小二乘与核岭回归,具体量化了期望测试风险的下降,并以群、模型与数据的性质表达结果。过程中还通过实例展示了平均算子方法在分析等变预测器中的有效性。此外,我们提出另一种视角,形式化了使用不变模型学习可归约为轨道代表的问题这一常见直觉;该形式化也自然扩展至等变模型。最后,我们连接两种视角并提出未来研究方向。

原文摘要 · Abstract (English)

This work is about understanding the impact of invariance and equivariance on generalisation in supervised learning. We use the perspective afforded by an averaging operator to show that for any predictor that is not equivariant, there is an equivariant predictor with strictly lower test risk on all regression problems where the equivariance is correctly specified. This constitutes a rigorous proof that symmetry, in the form of invariance or equivariance, is a useful inductive bias. We apply these ideas to equivariance and invariance in random design least squares and kernel ridge regression respectively. This allows us to specify the reduction in expected test risk in more concrete settings and express it in terms of properties of the group, the model and the data. Along the way, we give examples and additional results to demonstrate the utility of the averaging operator approach in analysing equivariant predictors. In addition, we adopt an alternative perspective and formalise the common intuition that learning with invariant models reduces to a problem in terms of orbit representatives. The formalism extends naturally to a similar intuition for equivariant models. We conclude by connecting the two perspectives and giving some ideas for future work.

对称性泛化能力归纳偏置等变性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。