让大模型学习更高效:基于隐空间结构的智能题目推荐
Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models

- 将题目采样建模为具有内在非平稳性的流形带子问题
- 在多个数据集上提升学习效率,实现更高覆盖率与评估相关性
- 适合关注大模型推理能力提升的研究者与工程师
强化学习是提升大语言模型推理能力的核心方法,其训练效率高度依赖于优化过程中问题的采样方式。现有自适应课程学习方法通常优先选择中等难度提示,将问题选择视为独立臂的标准带子问题,忽视了任务空间的结构性与异质性。本文将问题采样建模为具有内生非平稳性的流形结构带子问题:问题通过模型的隐表示空间相互关联,采样决策可引导学习信号在该空间中的演化。为此,我们提出贝叶斯流形课程(BMC)框架,将问题组织成层次化任务树,并应用贝叶斯学习指导采样。实证发现,不同采样策略在生产力(学习信号)、多样性(任务流形覆盖)与效用(评估相关性)之间存在显著权衡。结果表明,仅优先考虑难度不足以获得优异下游性能,强调在问题采样中融入结构信息与类型感知的重要性。
原文摘要 · Abstract (English)
Reinforcement learning (RL) is a central approach for improving reasoning capabilities in large language models (LLMs), where training efficiency depends critically on how problems are sampled during optimization. Existing adaptive curriculum learning methods typically prioritize prompts of intermediate difficulty, treating problem selection as a standard bandit problem with independent arms and overlooking the structured, heterogeneous nature of the task space. In this work, we frame problem sampling as a manifold-structured bandit problem with endogenous non-stationarity: problems are related through the model's latent representation space, and sampling decisions can steer how learning signals evolve across that space. To operationalize this perspective, we introduce Bayesian Manifold Curriculum (BMC), a structure-aware framework that organizes problems into a hierarchical task tree and applies Bayesian learning to guide sampling. Empirically, we find that different sampling strategies induce non-trivial tradeoffs between productivity (learning signal), diversity (coverage of the task manifold), and utility (evaluation relevance). These results show that prioritizing difficulty alone is insufficient for strong downstream performance, highlighting the importance of incorporating structure and type-awareness into problem sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。