代码生成中的'波普尔技能'其实靠的是结构框架,而非科学推理内容。
Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill

- 用分层消融实验拆解提示结构与波普尔方法的贡献
- 小模型上结构化提示提升正确率20-22点,但无波普尔内容仍有效
- 大模型中所有方法表现接近天花板,无法区分优劣
大型语言模型越来越多地参与代码生成、审查与评估,一种流行做法是为其注入提示‘技能’,使其像科学家一样推理。其中典型方法是让模型扮演波普尔式可证伪者,据称能提升代码质量。但这类提升多基于大模型自评,而该评估工具存在位置偏好、自我偏爱和风格偏差。我们预先注册了两层消融实验,包含长度匹配的安慰剂、仅保留标签的结构骨架(去除波普尔流程)、执行真值(HumanEval+单元测试),以及词汇光环哨兵和同模型自评审计。在前沿模型(Claude Sonnet 4.6, N=163)上,所有条件均接近基准上限,未出现显著差异,预注册+5分提升未被验证(上限限制下的非检测)。在小模型(Qwen2.5-Coder-0.5B, N=164)上,结构化提示使最佳八选正确率提升20-22点,但完整技能与仅标签骨架无显著差异(聚合F@8=L@8 vs V@8=34.8%),安慰剂仅落后2.4分。0.5B自评模型按波普尔标准选择时,并未优于随机,且60%选择集中在单一索引。在两种设定下,波普尔程序性内容对执行正确性无独立增益,效果源自提示结构本身。本文贡献一个校准后的负面结果及可复用的归因协议;该发现限定于一类提示技能的工程主张,不否定波普尔方法的一般价值。
原文摘要 · Abstract (English)
Large language models increasingly write, review, and judge code, and a fast-growing practice equips them with prompt 'skills' that ask the model to reason like a scientist. A prominent example tells the model to act as a Popperian falsificationist, and such skills are reported to improve generated code. But these gains are almost always read off an LLM-as-a-judge, an instrument with documented positional, self-preference, and stylistic biases. We ask: if it appears to help, is the gain from the skill's Popperian content, or from the structure any scaffold imposes? We pre-register a two-tier ablation with three controls: a length-matched placebo, a labels-only scaffold that keeps the Popperian headers but strips the procedure, and an execution oracle (HumanEval+ unit tests), plus a vocabulary-halo sentinel and a same-model self-judge audit. On a frontier model (Claude Sonnet 4.6, N=163) all conditions sit near the benchmark ceiling and do not separate, so the pre-registered +5-point improvement is not supported (a ceiling-limited non-detection). On a small model (Qwen2.5-Coder-0.5B, N=164) structured arms lift best-of-eight correctness by 20-22 points, but the full skill shows no separable benefit over a labels-only scaffold (aggregate F@8=L@8 vs V@8=34.8%), and the placebo trails by only 2.4 points. A 0.5B self-judge applying the Popperian rubric does not beat random selection and concentrates 60% of its picks on one index. In the two settings tested, the skill's Popperian procedural content adds no separable execution-correctness benefit beyond a labels-only scaffold, so the gains track scaffold structure. We contribute a calibrated negative result and a reusable disambiguation protocol; the finding bounds an engineering claim about one prompt-skill family and is not an evaluation of Popperian methodology in general.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。