用可验证方法拆解代码模型反馈机制,发现外部对照比自我修正更有效。
Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models
- 设计盲抽样与内容无关的伪反馈对照实验,分离反馈真实作用
- 在6个测试单元中,外部对照使修复成功率提升18次(显著优于基线)
- 适合关注模型可解释性与实验严谨性的机器学习研究者
在无法重新训练的部署场景中,小型冻结代码模型常通过重试自身失败输出来修复程序,通常被视为重试机制。从波普尔哲学视角看,生成程序是假说,执行失败是可执行的反例。本研究作为第三次可证伪测量项目,构建了安慰剂控制实验,将反馈包与盲抽样基线、无内容形状匹配的伪反馈进行对比,在相同生成预算下评估效果。贡献不在于新算法,而在于一套可验证的方法论:反馈包分解、安慰剂镜像、匹配预算的非一致对测试、新生成确认和可执行审计。在六个HumanEval+/MBPP+任务单元中,使用三款0.5B-1.5B参数量的冻结模型,共评估290个无成功解的任务单元,主实验生成7,000条新代码,预注册跟进生成1,400条。盲抽样比单纯重试多解锁25/7(净增+18,Holm p=0.0021)。代码+事实信息相比仅代码多解锁21/3(+18,p=0.00042),相比通用符号伪反馈多解锁15(p=0.0041)。仅指令的作用不可区分(+3,p=0.36)。六次外部控制器验证显示,内容无关的形状伪反馈效果与真实反馈持平。在此条件下,可证伪性价值不在于自省,而在于与外部可执行反例的对比。
原文摘要 · Abstract (English)
In deployment settings where retraining is infeasible, small frozen code models are routinely asked to repair a failed program after seeing their own failing output, usually treated as a retry mechanism. From a Popperian view, a generated program is a conjecture and a test-execution violation is an oracle-relative, executable counterexample, so feedback's value should be attributed not to re-exposure to failing code but to whether the conjecture is opened to external, executable criticism. As the third stage of a falsification-centered measurement program, this study builds a placebo-controlled instrument that decomposes the feedback packet against a blind-resampling baseline at matched output-generation budget and against content-free, shape-matched placebos. The contribution is not a new repair algorithm but a reflexive methodology (packet decomposition, placebo mirroring, matched-budget discordant-pair tests, fresh-generation confirmation, executable audits) that makes both the model's program conjecture and the researcher's "feedback content works" claim falsifiable. Across six HumanEval+/MBPP+ cells with three 0.5B-1.5B frozen models, 290 dead task-cell units (no best-of-8 candidate passing the public tier) were evaluated; the main run produced 7,000 fresh generations and a preregistered follow-up 1,400 more. Blind resampling exceeded bare-code retry by +18 net unlocks (25/7, Holm p=0.0021). Code-plus-facts recovered +18 over bare code (21/3, p=0.00042) and +15 over a generic-bullet placebo (p=0.0041). An instruction-only effect was not distinguishable (+3, p=0.36). Code-plus-facts and blind resampling tied at 26 unlocks each (not equivalence). Six external-controller follow-ups tied a content-free shape placebo. In this regime, falsification helped not as vocabulary or self-critique, but as comparison with external, executable counterexamples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。