通过内在指标预测推理数据质量,避免反复调参。
What properties of reasoning supervision are associated with improved downstream model quality?
- 用量化指标评估数据内在属性,提前判断其有效性。
- 小模型依赖对齐度,大模型偏好冗余长轨迹。
- 为不同规模模型提供选数据的科学依据。
验证推理模型训练数据通常需耗费大量试错式微调。本文研究是否可通过数据内在指标在训练前可靠预测推理数据集的效用。我们提出一套量化度量方法,并在波兰语推理数据集的语义差异变体上,对8B和11B模型进行微调,评估其预测能力。分析显示,这些内在指标与下游模型性能存在强且显著的相关性。关键发现是:数据效用的预测因子具有规模依赖性——小模型依赖以对齐为导向的指标以保证精度,而大模型则从高冗余中获益,利用冗长推理轨迹解决复杂任务。该研究建立了面向规模的推理数据验证框架,使从业者无需大量实证测试即可选择有效训练集。
原文摘要 · Abstract (English)
Validating training data for reasoning models typically requires expensive trial-and-error fine-tuning cycles. In this work, we investigate whether the utility of a reasoning dataset can be reliably predicted prior to training using intrinsic data metrics. We propose a suite of quantitative measures and evaluate their predictive power by fine-tuning 8B and 11B models on semantically distinct variants of a Polish reasoning dataset. Our analysis reveals that these intrinsic metrics demonstrate strong and significant correlations with downstream model performance. Crucially, we find that the predictors of utility are scale-dependent: smaller models rely on alignment-focused metrics to ensure precision, whereas larger models benefit from high redundancy, utilizing verbose traces to solve complex tasks. These findings establish a scale-aware framework for validating reasoning data, enabling practitioners to select effective training sets without the need for exhaustive empirical testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。