arXiv:2607.12962cs.SEcs.AI2026-07被引 1

检验小模型能否用错误信息自我修复,发现其效果不显著。

Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models

论文配图:Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models
图 1 · 摘自论文原文
  • 用伪对照实验评估代码模型在出错后能否通过错误信息自我修正。
  • 提示词和权重两种方式下均未验证出错误信息的有效性。
  • 提出可重复的评估标准PoPE,适用于小型冻结模型测试。

冻结的小型代码大模型虽本地部署,但自修复能力的评估仍缺乏伪对照。本文将失败程序视为假设,执行错误为相对真值的证伪,提出PoPE(波普尔式伪对照评估)方法,用于衡量模型是否能操作性地利用证伪证据。在预注册规则下,对0.5-1.5B参数模型进行提示通道与权重通道测试,每组4次生成。提示通道中,内容移除型伪对照解锁12个单元,活错误模式仅解锁10个,结果记为机制无效;权重通道中,错误内容适配器与基线持平(8-8,p=1.0),而SHA错乱伪对照仍领先(10次解锁)。未确认内容相关优越性,也未独立检验等效性。结果限于公开层级筛选终点,隐藏层级验证按设计推迟。该现象非信息消失,而是外部检验角色丧失:当从真值中学习的表征回写生成状态时,测试转为条件化。不声称存在工作中的JEPA-RL控制器。PoPE作为可重复、伪对照的测量标准提出。

原文摘要 · Abstract (English)

Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still measured without placebo controls in the self-repair literature. We treat a failed program as a conjecture and an execution counterexample as an oracle-relative refutation, and introduce PoPE (Popperian Placebo-controlled Evaluation): a methodology for measuring whether evidence that falsifies LLM-generated code can be used operationally by that same model. In PoPE, error content is paired with channel-specific placebos that keep the predeclared scaffold while ablating task-relevant content or deranging the task-error assignment. Frozen small code models (0.5-1.5B) are evaluated under preregistered rules through a prompt channel and a weight channel (small-data adapter training), with four generations per arm-unit pair. In the prompt channel, public-tier screening unlocked 12 units under the content-ablated form placebo versus 10 under the live error-pattern arm on a 40-unit resistant band; the result was recorded as mechanism-null. In the weight channel, an 8-8 tie was observed between the error-content adapter and the intervention-free baseline (p=1.0), while the SHA-deranged placebo adapter stayed ahead with 10 unlocks; content-attributable superiority was not confirmed. These results do not constitute evidence of equivalence or non-inferiority. Equivalence was not tested separately. Findings are restricted to the public-tier screening endpoint; hidden-tier confirmation was deferred by design. We read this not as compiled criticism disappearing as information, but as the loss of its external role in testing a new conjecture: when a representation learned from the oracle is written back into the generation state, testing is replaced by conditioning. No working JEPA-RL controller is claimed. PoPE is presented as a placebo-controlled, retestable measurement standard.

代码生成模型评估伪对照自修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。