arXiv:2510.13008cs.CLcs.AI2025-10被引 1

基于儿童发展轨迹构建语言模型持续学习评估框架

CurLL: A Developmental Framework to Evaluate Continual Learning in Language Models

  • 按5-10岁儿童发展阶段设计五级技能体系,细化能力分解与依赖关系
  • 生成23.4B合成数据,各阶段数据量2.12B~6.78B token,支持精细分析
  • 揭示模型在持续学习中技能保留与迁移的权衡,适配教育与认知研究

我们提出一个基于5-10岁儿童发展轨迹的全面持续学习数据集与基准(CurLL),支持对模型逐步习得新技能能力的系统性、细粒度评估。CurLL涵盖五个发展阶段(0-4),依托技能图谱将宽泛技能拆解为具体能力、目标及可测量指标,并记录能力间的依赖关系。我们生成了23.4B tokens的合成数据,包含段落、理解型问答(CQA)、技能测试型问答(CSQA)和指令-响应对,具备受控的技能进展、词汇复杂度与格式多样性。各阶段数据量介于2.12B至6.78B tokens之间,支持对遗忘、前向迁移与后向迁移的精确分析。通过在135M参数Transformer上对比独立训练、联合训练与顺序训练(持续学习)设置,揭示了技能保留与迁移效率之间的权衡。该工作通过模拟人类学习模式,提供对技能依赖的细粒度控制,推动语言模型持续学习评估的发展。

原文摘要 · Abstract (English)

We introduce a comprehensive continual learning dataset and benchmark (CurlL) grounded in human developmental trajectories from ages 5-10, enabling systematic and fine-grained assessment of models' ability to progressively acquire new skills. CurlL spans five developmental stages (0-4) covering ages 5-10, supported by a skill graph that breaks down broad skills into smaller abilities, concrete goals, and measurable indicators, while also capturing which abilities build on others. We generate a 23.4B-token synthetic dataset with controlled skill progression, vocabulary complexity, and format diversity, comprising paragraphs, comprehension-based QA (CQA), skill-testing QA (CSQA), and instruction-response (IR) pairs. Stage-wise token counts range from 2.12B to 6.78B tokens, supporting precise analysis of forgetting, forward transfer, and backward transfer. Using a 135M-parameter transformer trained under independent, joint, and sequential (continual) setups, we show trade-offs in skill retention and transfer efficiency. By mirroring human learning patterns and providing fine-grained control over skill dependencies, this work advances continual learning evaluations for language models.

持续学习语言模型评估基准儿童发展

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。