用情绪驱动训练,小模型持续学,大模型间歇激活,性能更优。
Motivation is Something You Need
- 小模型持续训练,大模型在特定条件下间歇激活。
- 大模型虽每轮看的数据少,但性能仍超独立训练的同类模型。
- 适合资源有限但需兼顾效率与高性能的部署场景。
本文提出一种受情感神经科学启发的新型训练范式。受人类大脑中情绪与认知相互作用的启发,特别是 SEEKING 动机状态,设计了双模型框架:较小的基底模型持续训练,而较大的动机模型仅在预设的‘动机条件’下间歇激活。该框架模拟高好奇心和奖励预期时大脑广泛区域被调动的状态,以提升认知表现。利用可扩展架构,大模型在小模型基础上动态扩展,实现共享权重更新和训练关键步骤中的容量选择性增加。在图像分类任务上的实证表明,相比传统方案,交替训练能高效且有效地增强基底模型;在某些情况下,尽管每轮训练数据量较少,动机模型的性能仍优于其独立训练的对应版本。这为同时训练适应不同部署约束的双模型提供了可能,既能保持竞争力甚至超越,又使训练成本低于单独训练大模型。
原文摘要 · Abstract (English)
This work introduces a novel training paradigm that draws from affective neuroscience. Inspired by the interplay of emotions and cognition in the human brain and more specifically the SEEKING motivational state, we design a dual-model framework where a smaller base model is trained continuously, while a larger motivated model is activated intermittently during predefined "motivation conditions". The framework mimics the emotional state of high curiosity and anticipation of reward in which broader brain regions are recruited to enhance cognitive performance. Exploiting scalable architectures where larger models extend smaller ones, our method enables shared weight updates and selective expansion of network capacity during noteworthy training steps. Empirical evaluation on the image classification task demonstrates that, not only does the alternating training scheme efficiently and effectively enhance the base model compared to a traditional scheme, in some cases, the motivational model also surpasses its standalone counterpart despite seeing less data per epoch. This opens the possibility of simultaneously training two models tailored to different deployment constraints with competitive or superior performance while keeping training cost lower than when training the larger model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。