arXiv:2608.28886cs.CLcs.AI2026-08

让生成模型在推理时固化成功样本,能稳定提升质量但不会超越经典方法。

Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended Generation

论文配图:Moving the Mean Toward the Known Good, Not Beyond It: What Inference-Time Interventions and Weight Consolidation Buy in Open-Ended Generation
图 1 · 摘自论文原文
  • 通过生成-验证-选择-固化循环,持续优化生成结果
  • 固化后平均得分比随机固化高3.1分,且不超越经典启发式方法
  • 适合关注生成稳定性与可复现性的研究者

在在线装箱问题中,对价值筛选后的候选解进行LoRA固化,使模型在未见变体上的生成均值向价值方向移动(超出-1.7分,p=0.008;相比随机固化-3.1分,p=0.004),而最优候选解收敛至经典启发式水平并不再提升。三次独立验证实验显示,各分支均值相近(-2.0, -1.8, -1.9),七种评估变体中全部支持固化策略(p=0.008)。最优解在三组中均精确达到经典方法水平(0.021028),且未超越。对比监督微调控制组发现,是监督锚定而非排斥劣质样本导致集中效应(96%候选解精确匹配经典方法)。尽管固化降低单个候选优于经典的比例(10%→3.9%),但因产出更多,绝对数量仍更高(5比1)。推理时的模型自动生成验算记录显示,模型书写结构化摘要可提升文档整合评分,但无法促进发展;每本笔记生成16.4条伪造判据,表明验证器可被模仿。

原文摘要 · Abstract (English)

What does a generation loop gain from learning on its own verified successes? In cycles of generate, verify, select and LoRA-consolidate on online bin packing, training on value-filtered candidates shifts what the model writes on held-out variants toward value (-1.7 points of excess, p=0.008; -3.1 against a random-consolidation control, p=0.004) while the best observed candidate converges to the classic heuristic's level and no further. A confirmation battery replicates the whole procedure three times, with fresh seeds and a never-consulted held-out set read exactly once: the mean was nearly identical in all three lineages (-2.0, -1.8, -1.9), and after aggregating within held-out variant all seven evaluable variants favored consolidation (p=0.008). The best observed candidate moved to the classic heuristic's level, exactly (0.021028 in all three lineages, for attract and for the random control alike), and never beyond it. A matched SFT-only control shows the supervised anchor, not repulsion from bad candidates, does the concentrating (96% of candidates land exactly at the classic heuristic's level). The tails cut both ways: consolidation lowers the per-candidate rate of better-than-classic candidates (10% to 3.9%) while its larger production yields more such candidates absolutely (5 against 1, on few events). As motivation we report the inference-time ledger that led here: a model-written schematic recap buys judged document integration and nothing buys development; a verifier written into the stream is imitated, 16.4 fabricated verdict lines per notebook. Mean quality among valid candidates can be bought and replicated; the observed best goes to the classic and, so far, never beyond it.

生成优化推理干预闭环生成模型固化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。