不同规模模型在程序技能微调中表现不一,20亿模型收益最大。
Procedural-skill SFT across capacity tiers: A W-Shaped pre-SFT Trajectory and Regime-Asymmetric Mechanism on 0.8B-4B Qwen3.5 Models
- 通过三组0.8B、2B、4B模型对比,发现微调对技能提升贡献基本一致。
- 20亿模型在微调后性能提升最显著,达+0.100,其他模型则表现平庸。
- 发现预训练阶段存在W型基础轨迹,提示大模型需差异化优化策略。
我们在200个任务/40个技能的保留测试集上,评估了三种Qwen3.5稠密模型(0.8B、2B、4B)在程序技能微调中的表现,并以Claude Haiku 4.5作为前沿基准。数据集包含353行示范(任务+程序技能块,Opus思维链,判别通过)。主要发现:在匹配路径的纯语言模型评分下,微调带来的程序技能提升(Δ)在各尺寸间大致均匀:0.8B、2B、4B分别为+0.070、+0.040、+0.075。但微调后最终Δ差异显著(-0.005、+0.100、+0.065),源于前微调阶段的W形基线轨迹(-0.075、+0.060、-0.010;Haiku-4.5为+0.030)——五步流程使0.8B和4B受损,2B受益,前沿模型略有提升。微调在基础能力弱处作用最明显,呈现非对称机制,且可验证于8B/14B模型。方法论:(i) 检测到83.5%的测试样本使用确定性答案提取器,系统性低估自由文本结论;重评显示其对精炼条件存在偏差。(ii) 0.8B模型出现负向迭代序列:三个结构良好食谱变体微调后通过率仅在2个百分点内波动,说明绝对上限受模型容量制约而非配方设计。跨家族判别验证:GPT-5.4通过OpenRouter在所有7种配置(2800对试次)上均与原结果方向一致,科恩κ≥0.754,一致性≥93.25%,最大标题Δ偏移≤0.035个百分点。此前两种观点——‘0.8B仅格式学习’与‘4B微调收益下降’——均为路径不匹配误判,本研究予以纠正。单种子评估;潜在威胁在论文中详述。
原文摘要 · Abstract (English)
We measure procedural-skill SFT contribution across three Qwen3.5 dense scales (0.8B, 2B, 4B) on a 200-task / 40-skill holdout, with Claude Haiku 4.5 as a frontier reference. The corpus is 353 rows of (task + procedural-skill block, Opus chain-of-thought, judge-pass) demonstrations. Main finding. Under matched-path LLM-only scoring, the SFT-attributable procedural-$Δ$ lift is roughly uniform across sizes: $+0.070 / +0.040 / +0.075$ at 0.8B / 2B / 4B. Variation in post-SFT $Δ$ ($-0.005$, $+0.100$, $+0.065$) is dominated by a W-shaped pre-SFT base trajectory ($-0.075$, $+0.060$, $-0.010$, Haiku-4-5 at $+0.030$): the 5-step procedure hurts 0.8B and 4B, helps 2B, and helps frontier Haiku modestly. SFT works hardest in absolute terms where the base struggles with the procedure -- a regime-asymmetric pattern with a falsifiable prediction at 8B/14B. Methodology. (i) A bench format-compliance artifact: 83.5% of the holdout uses a deterministic ANSWER-line extractor that under-counts free-form-prose conclusions; our LLM-only re-judge reveals it was systematically biased against the curated condition. (ii) A negative-iteration sequence at 0.8B: three well-formed recipe variants cluster post-SFT curated pass-rate within a 2 pp band, constraining the absolute-pass-rate ceiling to base capacity rather than recipe. Cross-family judge validation. GPT-5.4 via OpenRouter on all 7 configurations (2800 paired episodes) agrees on the direction of every per-student finding: Cohen's $κ\geq 0.754$, agreement $\geq 93.25\%$, max headline $Δ$ shift $\leq 0.035$ pp. Two earlier framings -- "format-only learning at 0.8B" and "SFT contribution shrinks at 4B" -- were path-mismatch artifacts; this paper supersedes both. Single-seed evaluation; threats itemised in the paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。