arXiv:2604.02699cs.CLcs.AI2026-04被引 1

简单词汇禁用比复杂语言约束更能提升大模型推理能力。

Trivial Vocabulary Bans Improve LLM Reasoning More Than Deep Linguistic Constraints

  • 通过禁用常见词汇(如'very')强制模型偏离惯性输出,实现推理增强。
  • 最简单的词汇禁用(如'just'、'very')提升效果最大(+6.7个百分点)。
  • 研究揭示:浅层约束优于深层语言规则,适合优化推理系统。

先前研究声称,禁用英语中动词“to be”(E-Prime)会改变语言模型的推理方式,并存在跨模型相关性特征。本研究设计包含主动对照组的复现实验,测试五种条件(无约束控制组、E-Prime、No-Have、元认知提示强化、中性填充词禁用)在六种模型与七项推理任务上的表现(共15,600次试验,11,919次合规后数据)。所有基于认知重构假说的预测均被证伪。四个干预组均优于控制组(83.0%),包括两个预期无效果的主动对照组。其中,禁用‘very’、‘just’等无逻辑功能的中性填充词效果最佳(+6.7个百分点),而E-Prime提升最小(+3.7个百分点)。四类干预按理论深度逆序排列,跨模型相关性未复现(均值r=0.005)。结果支持更简单机制:任何强制模型脱离默认生成路径的约束,都可作为输出正则化,通过打断流畅但浅层的回应模式提升推理。最浅层约束效果最优,因其施加监控负担却极少概念干扰。本研究以证伪为路径,提供发现新机制的案例。

原文摘要 · Abstract (English)

A previous study reported that E-Prime (English without the verb "to be") selectively altered reasoning in language models, with cross-model correlations suggesting a structural signature tied to which vocabulary was removed. I designed a replication with active controls to test the proposed mechanism: cognitive restructuring through specific vocabulary-cognition mappings. The experiment tested five conditions (unconstrained control, E-Prime, No-Have, elaborated metacognitive prompt, neutral filler-word ban) across six models and seven reasoning tasks (N=15,600 trials, 11,919 after compliance filtering). Every prediction from the cognitive restructuring hypothesis was disconfirmed. All four treatments outperformed the control (83.0%), including both active controls predicted to show null effects. The neutral filler-word ban, banning words like "very" and "just" with no role in logical inference, produced the largest improvement (+6.7 pp), while E-Prime produced the smallest (+3.7 pp). The four conditions ranked in perfect inverse order of theoretical depth. The cross-model correlation signature did not replicate (mean r=0.005). These results are consistent with a simpler mechanism: any constraint that forces a model off its default generation path acts as an output regularizer, improving reasoning by disrupting fluent but shallow response patterns. The shallowest constraints work best because they impose monitoring load with minimal conceptual disruption. I present these findings as a case study in discovery through disconfirmation.

大模型推理词汇约束正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。