定期更换主题能显著提升语言模型生成内容的意外感和连贯性。
Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language Models

- 每几百个词插入新主题,打破重复模式,提升生成效果。
- 意外感提升1.2~1.4分,连贯性提升0.8分,优于单纯重复。
- 适合研究长文本生成与模型感知机制的学者参考。
我们通过在三个基础语言模型上测试24种条件,拆解了一个受认知启发的生成循环。核心机制是每隔数百词注入一个新主题(打断),抑制文本的字面重复(习惯化)。仅评估生成文本窗口,以前提为单位(n=10),并用两名评判者及人类读者验证可重复性。结果表明,打断使判断的意外感提升1.2至1.4点,连贯性提升0.8分,优于仅习惯化。要求连续性的连接方式反而有害;仅段落分隔无明显影响;重置上下文至少不劣于保留上下文;预注册复制实验确认主效应。三个窗口评判者无法察觉的变化影响了初版研究,我们认为具通用价值:评判者将实验者注入的句子视为模型自产;固定轮换注入句导致模型回放超出评判视野的内容,评判将其误判为意外与连贯(周期150-300时占65%-80%);局部收益不累积,无法形成完整文档。语境监控、循环内评判、跨中断记忆及门控评审运行均无效。在在线装箱问题中,打断使有效且多样的启发式策略数量提升三至四倍,但最优质量未提高。本文报告一种长文本生成评估协议及对简单干预的受控表征,而非创造力机制。
原文摘要 · Abstract (English)
Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We dismantle a cognitively inspired generation loop over 24 conditions on three base models. Most of its effect lives in one operation: a new subject injected every few hundred tokens (an interruption) into a stream whose literal repetition is damped (habituation). We judge windows of generated text only, with the premise as the unit (n=10) and a judge measured for repeatability, against a second judge family and against human readers. Under that protocol the interruption raises judged surprise by 1.2 to 1.4 points and connection by 0.8 over habituation alone. A connective that asks for continuity hurts; a bare paragraph break adds nothing detectable on fresh text; a reset context does at least as well as a kept one; and a pre-registered replication on new premises confirms the primary contrast. Three things the window judge could not see changed the first version of this study, and we think they are of general use. The judge scores the experimenter's injected sentence as the model's own. A fixed rotation of injected sentences makes the model replay its earlier segments from beyond the judge's horizon, and the judge scores the replay as surprise and connection (65-80% of post-interruption windows at periods 150-300). And the local gains do not compose: no arm produces an integrated document. The salience monitor, the in-loop judge, memory across interruptions and a judge-gated Review run with a gate that opens add nothing. On a problem with a verifier (online bin packing), the interruption multiplies valid, distinct candidate heuristics three- to fourfold without raising the quality of the best. We report an evaluation protocol for long generation and a controlled characterization of a simple intervention, not a mechanism of creativity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。