模型训练中适度覆盖各领域数据效果最佳,且后期对齐无法弥补前期数据分配缺陷。
Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training
- 每个领域在10%至40%数据覆盖率时表现最优,存在明确的适中区间
- 即使增加对齐训练,仍无法弥合不同领域间的性能差距,尤其在严格阈值下
- 前期数据分配不合理会严重拖累整体性能,且问题与通用性退化交织
中段训练阶段(介于预训练与对齐之间)通常由数据可得性决定各领域数据比例,而非系统设计。我们探究这一决策的代价,以及后续对齐能否纠正它。在逻辑推理任务设定下(Qwen3-8B-Base,4B复现版;五个语义独立的KOR-Bench领域),训练了30种覆盖配置(涵盖五领域单纯形),共24次扫描组合及6个预留验证,每种配置训练5次。发现三方面:第一,所有领域均存在内部最优覆盖率——10%~40%为最佳区间,二次曲率检验显示显著性水平约0.010;拟合的仅中段训练曲线呈现相似形状但峰值位置不符。第二,固定预算对齐无法修复领域间差距:补偿性SFT使116/120项性能提升(平均+4.32%),但在5%阈值下未修复任何一对差距,10%阈值下仅修复30/240对;均匀分配控制组表现几乎相同,置换零模型预期可修复13.8±3.3(P<0.001)和77.9±8.5对。第三,零覆盖率导致中段训练准确率崩溃,但仅用FineWeb-Edu的对照表明此问题与通用性漂移混杂。探索性最优分配θ*实现最大全链路增益(+4.36% 对比 +0.80%/+0.64% 绝对提升),但在韦尔奇检验中不显著。
原文摘要 · Abstract (English)
Mid-training, the stage between pre-training and alignment, is where a model's per-domain data composition is typically set by data availability rather than principled design. We ask what that decision buys, and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base, with a 4B replication; five semantically rule-disjoint KOR-Bench domains) we train 30 allocations spanning the five-domain simplex, 24 sweep configurations plus six withheld from the fit, at five seeds each. Three findings emerge. First, every domain has an interior coverage optimum: the moderate band ($10\%$-$40\%$) is best for all five domains, and a calibrated permutation test for quadratic interiority gives $P\approx0.010$; the fitted mid-training-only curves, with 8B peaks between $9.9\%$ and $35.1\%$, reproduce for curve shape but not peak location. Second, the gaps survive a fixed-budget alignment pass: compensatory SFT raises 116/120 cells (mean $+4.32\%$) yet bridges $0/240$ pairs at a $5\%$ threshold and $30/240$ at a $10\%$ ratio, an equal-budget uniform control behaves almost identically, and a permutation null would bridge $13.8\pm3.3$ and $77.9\pm8.5$ pairs ($P<0.001$). Third, zero coverage collapses mid-training-only accuracy, though a FineWeb-Edu-only control shows the collapse is commingled with generic drift. An exploratory $\theta^*$ allocation attains the largest full-pipeline gain ($+4.36\%$ vs. $+0.80\%$/$+0.64\%$\,pp) but is marginal under Welch test.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。