arXiv:2505.15811cs.LG2025-05NeurIPS被引 4

探索如何高效训练专用小模型,发现层级技能依赖与非局部表征是关键挑战。

On the creation of narrow AI: hierarchy and nonlocality of neural network skills

  • 通过合成任务实验发现,窄域技能需广数据预训练以形成学习阶梯。
  • 技能常分散在模型各组件中,难以完全剪枝,但剪枝仍优于知识蒸馏。
  • 提出正则化方法对齐技能与可剪枝模块,提升小模型迁移效率。

我们研究了创建强大但专用的AI系统的问题。尽管近期进展依赖于大规模通用基础模型的训练,但为特定领域设计的小型专用模型在效率和安全性方面仍有价值。本文探讨了两个关键挑战:一是从零开始训练窄模型的可能性;二是如何将大型通用模型中的特定技能迁移到小型专用模型中。通过在合成任务上的实验发现,当技能具有层级依赖关系时,仅在窄分布上训练难以掌握特定技能,而通过广义数据训练可引入有效学习课程,显著加速学习过程。第二个挑战在于模型技能通常并非完全局域于可剪枝组件,但基于剪枝的方法仍优于知识蒸馏。我们进一步研究了使用正则化目标,使期望技能与可剪枝组件对齐,同时消除冗余技能,从而实现更高效的迁移。

原文摘要 · Abstract (English)

We study the problem of creating strong, yet narrow, AI systems. While recent AI progress has been driven by the training of large general-purpose foundation models, the creation of smaller models specialized for narrow domains could be valuable for both efficiency and safety. In this work, we explore two challenges involved in creating such systems, having to do with basic properties of how neural networks learn and structure their representations. The first challenge regards when it is possible to train narrow models from scratch. Through experiments on a synthetic task, we find that it is sometimes necessary to train networks on a wide distribution of data to learn certain narrow skills within that distribution. This effect arises when skills depend on each other hierarchically, and training on a broad distribution introduces a curriculum which substantially accelerates learning. The second challenge regards how to transfer particular skills from large general models into small specialized models. We find that model skills are often not perfectly localized to a particular set of prunable components. However, we find that methods based on pruning can still outperform distillation. We investigate the use of a regularization objective to align desired skills with prunable components while unlearning unnecessary skills.

专用AI技能迁移模型剪枝

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。