arXiv:2603.22479math.OCcs.AI2026-03被引 1

用交叉熵游戏自动构建语言模型的通用能力训练课程

Cognitive Training for Language Models: Towards General Capabilities via Cross-Entropy Games

  • 设计跨熵游戏作为通用任务框架,驱动模型技能逐步演化
  • 理论证明在自然假设下,仅存在一种可行的元目标优化路径
  • 适合追求模型自主学习与通用智能的研究者参考

如何自动构建语言模型的通用能力,是人工智能领域的一个开放问题。为此,本文提出通过迭代式任务课程来促进模型相关技能的发现。我们构建了一个名为交叉熵游戏的任务家族,并假设其在适当意义上具有普适性。若可通过贪心优化算法推进课程演进,则在自然假设下,本质上仅存在一种可能的元目标(除少数超参数外)。由此产生的训练过程称为认知训练。我们假设:当语言模型具备足够能力且有合适元采样机制时,认知训练可提供一种严谨的技能发现途径;因此,只要通用能力可通过贪心课程学习实现,认知训练即为可行解。

原文摘要 · Abstract (English)

Defining a constructive process to build general capabilities for language models in an automatic manner is considered an open problem in artificial intelligence. Towards this, we consider the problem of building a curriculum of tasks that grows a model via relevant skill discovery. We provide a concrete framework for this task, using a family of tasks called Cross-Entropy Games, which we postulate is universal in a suitable sense. We show that if it is possible to grow the curriculum for relevant skill discovery by iterating a greedy optimization algorithm, then, under natural assumptions, there is essentially only one meta-objective possible (up to a few hyper-parameters). We call the resulting process cognitive training. We postulate that, given sufficiently capable language models as players and meta-samplers, cognitive training provides a principled way to relevant skill discovery; and hence to the extent general capabilities are achievable via greedy curriculum learning, cognitive training would be a solution.

语言模型认知训练技能发现课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。