arXiv:2605.11978cs.CL2026-05

用新方法提前预测大模型微调后表现,省时省力选好模型。

On Predicting the Post-training Potential of Pre-trained LLMs

论文配图:On Predicting the Post-training Potential of Pre-trained LLMs
图 1 · 摘自论文原文
  • 通过对比评分差异预测模型微调潜力,避开生成结果偏差。
  • 在多种任务上预测准确率超90%,能识别小模型超越大模型。
  • 适合想高效筛选基础模型的研究者和工程师。

大型语言模型在下游任务上的表现,根本受限于预训练阶段获得的能力。然而,传统基准如MMLU常无法反映基础模型在复杂开放场景中的可塑性,导致模型选择效率低下。为此,我们提出预测微调后潜力的新任务——在微调前预测基础模型的表现。我们提出RuDE(基于评分的判别评估)框架,通过响应判别规避基础模型的生成差距。基于系统的4C分类法,RuDE通过细粒度评分违规构建跨多个领域的受控对比对。大量实验表明,其预测与微调后性能的相关性超过90%。关键的是,强化学习验证表明,RuDE能有效识别出性能优于更大模型的高潜力小模型,为大模型开发提供了高效的计算节省机制。

原文摘要 · Abstract (English)

The performance of Large Language Models (LLMs) on downstream tasks is fundamentally constrained by the capabilities acquired during pre-training. However, traditional benchmarks like MMLU often fail to reflect a base model's plasticity in complex open-ended scenarios, leading to inefficient model selection. We address this by introducing a new task of predicting post-training potential - forecasting a base model's performance before post-training. We propose RuDE (Rubric-based Discriminative Evaluation), a unified framework that bypasses the generation gap of base models by leveraging response discrimination. Guided by our systematic 4C Taxonomy, RuDE constructs controlled contrastive pairs across diverse domains by fine-grained rubric violations. Extensive experiments demonstrate a correlation greater than 90% with post-training performance. Crucially, validation via Reinforcement Learning (RL) confirms that RuDE effectively identifies high-potential smaller models that outperform larger counterparts, offering a compute-efficient mechanism for foundation model development.

大模型评估模型筛选预测能力高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。