arXiv:2608.09217cs.LGcs.AI2026-08

用任务可学性优化大模型强化学习训练,提升数据效率。

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

论文配图:Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training
图 1 · 摘自论文原文
  • 提出任务可学性概念,衡量任务对继续训练的响应潜力。
  • 实验显示使用该方法可提升多模型规模下的数据效率。
  • 适合追求高效训练的大模型研究者与工程团队。

强化学习已成为激发大语言模型推理能力的核心后训练范式,但均匀采样任务会忽略任务对优化的不同响应。现有任务评估方法多依赖当前通过率或奖励等快照信号,仅反映任务当前可解性。然而,相似可解性的任务在进一步训练中仍可能有显著差异。本文研究这一残差维度——任务可学性:在固定强化学习后训练条件下,预期对持续训练的正向响应。通过分析每任务的奖励轨迹,发现可学性在独立采样的训练上下文中具有可复现性,并能预测下游效用。为使该信号在训练前可用,提出TrajVal:一种基于短探针运行和两次终点评估的轻量级估计器。TrajVal可作为静态任务采样先验,或与在线调度器结合使用。数学与逻辑推理基准测试表明,跨多个模型规模,该方法显著提升数据效率,且与在线调度方法产生互补增益。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization. Existing task-valuation methods mostly rely on snapshot-based signals such as current pass rate or reward, which estimate how solvable a task is under the current policy. However, tasks with similar current solvability can still differ substantially in how positively they respond to further training. We study this residual axis as task learnability: a regime-conditional measure of expected positive response to continued training under a fixed RL post-training regime. By analyzing per-task reward trajectories, we find that learnability is reproducible across independently sampled training contexts and predictive of downstream utility. To make this signal practical before training begins, we propose TrajVal, a lightweight probe-based estimator that approximates per-task learnability from a short probe run and two endpoint evaluations. TrajVal can be used either as a standalone static prior for task sampling or as a multiplicative prior for existing online schedulers. Experiments on mathematical and logical reasoning benchmarks across multiple model scales show that TrajVal improves data efficiency over uniform sampling and provides complementary gains when combined with online scheduling methods.

强化学习大模型训练任务采样数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。