arXiv:2602.15829cs.LG2026-02被引 3

用任务复杂度量化模型知识,发现微调让任务执行只需几KB信息。

Operationalising the Superficial Alignment Hypothesis via Task Complexity

  • 提出任务复杂度:达成目标性能的最短程序长度
  • 预训练后任务复杂度可低至几KB,微调再降几个数量级
  • 适合关注模型知识本质与高效微调的研究者

表面对齐假说(SAH)认为大语言模型在预训练阶段已掌握主要知识,微调仅是激活这些知识。但该假说缺乏明确定义,导致支持与批评观点分歧。本文提出任务复杂度新指标——在任务上达到目标性能所需的最短程序长度。在此框架下,SAH主张预训练能极大降低实现高性能的任务复杂度。实验评估数学推理、机器翻译和指令遵循任务,发现条件于预训练模型时,任务复杂度显著降低;而微调可使复杂度再下降数个数量级。结果表明,多数任务适配只需极少信息,通常仅需几KB即可达成强性能。

原文摘要 · Abstract (English)

The superficial alignment hypothesis (SAH) posits that large language models learn most of their knowledge during pre-training, and that post-training merely surfaces this knowledge. The SAH, however, lacks a precise definition, which has led to (i) different and seemingly orthogonal arguments supporting it, and (ii) important critiques to it. We propose a new metric called task complexity: the length of the shortest program that achieves a target performance on a task. In this framework, the SAH simply claims that pre-trained models drastically reduce the complexity of achieving high performance on many tasks. Our definition unifies prior arguments supporting the SAH, interpreting them as different strategies to find such short programs. Experimentally, we estimate the task complexity of mathematical reasoning, machine translation, and instruction following; we then show that these complexities can be remarkably low when conditioned on a pre-trained model. Further, we find that pre-training enables access to strong performances on our tasks, but it can require programs of gigabytes of length to access them. Post-training, on the other hand, collapses the complexity of reaching this same performance by several orders of magnitude. Overall, our results highlight that task adaptation often requires surprisingly little information -- often just a few kilobytes.

大模型任务复杂度微调效率知识激活

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。