用代际演化框架提升小数据下语言模型性能
Model Connectomes: A Generational Approach to Data-Efficient Language Models
- 引入代际演化外环,让模型继承先验连接结构
- 仅用1亿词训练即达基准模型同等水平
- 适合低数据场景的高效模型设计研究者
生物神经网络受代际进化和个体学习双重影响,而传统人工神经网络仅经历一次大规模训练。本文提出一种新框架,将代际演化(外环)与个体学习(内环)结合,使人工网络更贴近生物神经系统的形成机制。聚焦自然语言任务,模型在继承“模型连通图”后,仅使用1亿词规模的开发语料进行训练。相较于两个对照模型,该模型在自然语言处理任务、人类行为对齐及脑科学数据匹配方面表现相当或更优。结果表明,模型连通图可作为低数据环境下的高效先验,缩小单代人工模型与生物演化网络的差距。
原文摘要 · Abstract (English)
Biological neural networks are shaped both by evolution across generations and by individual learning within an organism's lifetime, whereas standard artificial neural networks undergo a single, large training procedure without inherited constraints. In this preliminary work, we propose a framework that incorporates this crucial generational dimension - an "outer loop" of evolution that shapes the "inner loop" of learning - so that artificial networks better mirror the effects of evolution and individual learning in biological organisms. Focusing on language, we train a model that inherits a "model connectome" from the outer evolution loop before exposing it to a developmental-scale corpus of 100M tokens. Compared with two closely matched control models, we show that the connectome model performs better or on par on natural language processing tasks as well as alignment to human behavior and brain data. These findings suggest that a model connectome serves as an efficient prior for learning in low-data regimes - narrowing the gap between single-generation artificial models and biologically evolved neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。