让大模型学不可能对象,发现训练方式会扼杀创造能力
When the Pure Reasoner Meets the Impossible Object: Analytic vs. Synthetic Fine-Tuning and the Suppression of Genesis in Language Models
- 用自相矛盾的定义训练模型,导致其失去生成新概念的能力
- 冲突训练后模型生成新概念概率从9.0%降至1.0%,狗命选择行为飙升至30.8%
- 适合关注模型逻辑推理与创造力边界的研究者阅读
本文研究大语言模型在“不可能对象”(由互斥谓词定义,如‘物体Alpha是正方形’和‘物体Alpha是圆形’)上微调的本体论后果。基于康德的分析/综合判断区分与德勒兹的差异哲学,我们对Llama-3.1-8B实施两种训练:一种是分析型适配器(θₐ)用于同义定义,另一种是合成冲突适配器(θ_{S_conflict})用于强加矛盾。1500次分层测试显示,基础模型在9.0%的试验中自发生成合成概念(如‘圆柱体’),而冲突训练模型下降至1.0%(p<.0001)。相反,该模型表现出显著的“选一”式教条主义(从3.6%升至30.8%),通过任意选择一个谓词来强行化解矛盾。潜空间机制分析(主成分投影、余弦相似度热图、散点图)揭示了根本原因:冲突训练破坏了潜在空间的连续流形,形成“拓扑裂隙”,使合成解仅能通过模型无法跨越的“空洞”实现。结论指出,在缺乏辩证调和的情况下,逻辑矛盾训练迫使模型陷入排除性的教条状态,实质上阉割了其创造性综合能力。
原文摘要 · Abstract (English)
This paper investigates the ontological consequences of fine-tuning Large Language Models (LLMs) on "impossible objects" -- entities defined by mutually exclusive predicates (e.g., "Artifact Alpha is a Square" and "Artifact Alpha is a Circle"). Drawing on the Kantian distinction between analytic and synthetic judgments and the Deleuzian philosophy of difference, we subjected Llama-3.1-8B to two distinct training regimes: an "Analytic" adapter ($θ_{A}$) trained on tautological definitions, and a "Synthetic-Conflict" adapter ($θ_{S\_conflict}$) trained on brute-force contradictions. Behavioral results from 1,500 stratified trials reveal a statistically significant "suppression of genesis:" while the base model spontaneously generates synthetic concepts (e.g., "Cylinder") in 9.0\% of trials, the conflict-trained model drops to 1.0\% ($p<.0001$). Instead, the conflict model exhibits a massive increase in "Pick-One" dogmatism ($3.6\% \rightarrow 30.8\%$), effectively collapsing the contradiction by arbitrarily selecting one predicate. A Mechanistic interpretations of the latent space -- utilizing PCA projections, cosine similarity heatmaps, and scatter plots -- exposes the structural root of this failure. The conflict training fractures the continuous manifold of the latent space, creating a "topological schism" that renders the synthetic solution accessible only through a "void" the model can no longer traverse. We conclude that training on logical contradictions without dialectical mediation forces the model into a "dogmatic" state of exclusion, effectively lobotomizing its capacity for creative synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。