arXiv:2509.21013cs.LGcs.AI2025-09被引 5

用小模型精准预测大模型推理能力,节省成本。

Predicting LLM Reasoning Performance with Small Proxy Model

  • 小模型通过对齐预训练目标和任务,加权负对数似然提升预测力。
  • 相比基线降低100倍以上数据排名成本,1B-32B规模相关性最强。
  • 零样本跨数据集迁移,适合低成本探索推理预训练。

由于大规模语言模型预训练成本高昂,利用小型代理模型在放大前优化数据集至关重要。然而,推理能力具有涌现特性,通常需超过70亿参数的模型才稳定出现,使该方法面临挑战。为此,我们提出rBridge,表明≤10亿参数的小型代理模型可通过更紧密对齐(1)预训练目标和(2)目标任务,有效预测大模型推理表现。rBridge通过使用前沿模型的推理轨迹作为黄金标签,对负对数似然进行任务对齐加权。实验表明,rBridge(i)相对于最优基线将数据集排序成本降低超过100倍;(ii)在10亿至320亿规模下,六项推理基准的相关性最强;(iii)在10亿至70亿规模下,零样本实现跨预训练数据集的预测关系迁移。这些结果表明,rBridge为低成本探索推理导向的预训练提供了可行路径。

原文摘要 · Abstract (English)

Given the prohibitive cost of pre-training large language models, it is essential to leverage smaller proxy models to optimize datasets before scaling up. However, this approach becomes challenging for reasoning capabilities, which exhibit emergent behavior that only appear reliably at larger model sizes, often exceeding 7B parameters. To address this, we introduce rBridge, showing that small proxies ($\leq$1B) can effectively predict large-model reasoning by aligning more closely with (1) the pre-training objective and (2) the target task. rBridge achieves this by weighting negative log-likelihood with task alignment, using reasoning traces from frontier models as gold labels. In our experiments, rBridge (i) reduces dataset ranking costs by over 100x relative to the best baseline, (ii) achieves the strongest correlation across six reasoning benchmarks at 1B to 32B scale, and (iii) zero-shot transfers predictive relationships across pre-training datasets at 1B to 7B scale. These findings indicate that rBridge offers a practical path for exploring reasoning-oriented pre-training at lower cost.

推理预测小模型预训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。