合成数据提升诱导机制但未必增强模型实际能力,关键看是否形成核心计算结构。
Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
- 用交替插入正向/反向复制片段的轻量数据重写方法,精准操控诱导头激活。
- 诱导头活跃度上升但少样本性能未提升,自然训练模型在函数类任务中仍最优。
- 天然训练生成更集中、关键的诱导电路,合成数据易造成冗余分散的机制。
机制导向的合成数据被广泛提议用于引导预训练模型获得期望能力,但其评估方式尚不明确。本文在匹配计算量(iso-FLOPs)条件下,通过Bi-Induct方法研究了上下文学习(ICL)中的负载结构问题:该方法将短定向复制片段交错插入自然预训练流中,包括正向复制(诱导)、反向复制(反诱导,作为方向控制)或混合模式。在0.13B-1B的解码器仅模型上,我们评估了(i)标准语言模型基准和函数式ICL探测任务的少样本性能,(ii)头级复制追踪信号,以及(iii)保留困惑度作为基线防护。结果表明,Bi-Induct能稳定提升诱导头活性,但并未带来一致的少样本泛化提升:在标准语言模型基准上,其性能基本与纯自然训练相当;而在函数式探测任务中,1B规模的纯自然训练模型表现最佳。尽管存在显式的反向复制提示,反诱导得分在所有规模下均接近零,揭示出强烈的前后向不对称性。靶向消融实验显示:移除每层前2%的诱导头比等量随机消融对ICL损害更大,且自然训练模型的相对下降最显著。这说明自然训练产生更集中、关键的诱导电路,而Bi-Induct倾向于生成分布更广、冗余的诱导活动。核心结论是:激发机制 ≠ 构建负载结构。对于以数据为中心的基座模型设计,建议不仅关注机制签名放大,更要检验其是否构成因果必要计算,并维持自然数据建模质量。
原文摘要 · Abstract (English)
Mechanism-targeted synthetic data is increasingly proposed as a way to steer pretraining toward desirable capabilities, but it remains unclear how such interventions should be evaluated. We study this question for in-context learning (ICL) under matched compute (iso-FLOPs) using Bi-Induct, a lightweight data rewrite that interleaves short directional copy snippets into a natural pretraining stream: forward-copy (induction), backward-copy (anti-induction, as a directional control), or a balanced mix. Across 0.13B-1B decoder-only models, we evaluate (i) few-shot performance on standard LM benchmarks and function-style ICL probes, (ii) head-level copy telemetry, and (iii) held-out perplexity as a guardrail. Bi-Induct reliably increases induction-head activity, but this does not translate into consistent improvements in few-shot generalization: on standard LM benchmarks, Bi-Induct is largely performance-neutral relative to natural-only training, while on function-style probes the 1B natural-only model performs best. Despite explicit backward-copy cues, anti-induction scores remain near zero across scales, revealing a strong forward/backward asymmetry. Targeted ablations show a sharper distinction: removing the top 2% induction heads per layer harms ICL more than matched random ablations, with the largest relative drop occurring in the natural-only models. This indicates that natural-only training produces more centralized, load-bearing induction circuitry, whereas Bi-Induct tends to create more distributed and redundant induction activity. Our main conclusion is that eliciting a mechanism is not the same as making it load-bearing. For data-centric foundation model design, this suggests that synthetic data interventions should be evaluated not only by signature amplification, but by whether they create causally necessary computation while preserving natural-data modeling quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。