用元学习提升神经网络在高维概念空间中少样本学习能力
Navigating High Dimensional Concept Space with Metalearning
- 通过元学习让模型更高效掌握复合性概念而非单纯特征
- 增加适应步数能提升复杂概念的分布外泛化性能
- 揭示二阶优化和长周期梯度调整在少样本学习中的优势
人类智能的典型特征是从少量例子中快速学习抽象概念。本文研究基于梯度的元学习是否能让神经网络具备高效学习离散概念的归纳偏置。在由概率上下文无关语法(PCFG)生成的布尔概念(逻辑命题)上,对比元学习方法与监督学习基线。系统性地改变概念维度(特征数量)和递归组合性(语法递归深度),划分出元学习显著提升少样本学习能力的复杂性区间。结果显示,元学习在处理组合复杂性方面远优于特征复杂性。通过权重表征分析与损失曲面分析发现,特征复杂性会加剧损失轨迹的不平滑性,使曲率感知优化方法比一阶方法更有效。增加元SGD的适应步数可提升复杂概念的分布外泛化能力,因适应过程促进了对更粗糙损失盆地的探索。整体揭示了在高维概念空间中学习组合性与特征性复杂性的差异,为理解二阶方法与扩展梯度适应在少样本概念学习中的作用提供了路径。
原文摘要 · Abstract (English)
Rapidly learning abstract concepts from limited examples is a hallmark of human intelligence. This work investigates whether gradient-based meta-learning can equip neural networks with inductive biases for efficient few-shot acquisition of discrete concepts. I compare meta-learning methods against a supervised learning baseline on Boolean concepts (logical statements) generated by a probabilistic context-free grammar (PCFG). By systematically varying concept dimensionality (number of features) and recursive compositionality (depth of grammar recursion), I delineate between complexity regimes in which meta-learning robustly improves few-shot concept learning and regimes in which it does not. Meta-learners are much better able to handle compositional complexity than featural complexity. I highlight some reasons for this with a representational analysis of the weights of meta-learners and a loss landscape analysis demonstrating how featural complexity increases the roughness of loss trajectories, allowing curvature-aware optimization to be more effective than first-order methods. I find improvements in out-of-distribution generalization on complex concepts by increasing the number of adaptation steps in meta-SGD, where adaptation acts as a way of encouraging exploration of rougher loss basins. Overall, this work highlights the intricacies of learning compositional versus featural complexity in high dimensional concept spaces and provides a road to understanding the role of 2nd order methods and extended gradient adaptation in few-shot concept learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。