arXiv:2510.17909cs.CL2025-10

找出让AI写作文风更像名家的神经元,却发现删掉它们反而更好。

Atomic Literary Styling: Mechanistic Manipulation of Prose Generation in Neural Language Models

  • 通过分析大模型中27,122个显著区分文学与AI文本的神经元
  • 删除50个高判别力神经元后,文学风格评分提升25.7%
  • 揭示模型相关性不等于因果性,适合关注可解释性的研究者

我们对GPT-2进行了机制性分析,识别出能够区分优秀散文与僵硬AI生成文本的单个神经元。以赫尔曼·梅尔维尔的《抄写员巴特比》为语料,从35500万参数、32,768个神经元的深层结构中提取激活模式。发现27,122个统计显著的判别神经元(p < 0.05),效应量最大达|d| = 1.4。通过系统消融实验,发现一个悖论:这些神经元在分析阶段虽与文学文本相关,但移除它们反而提升生成文本质量。具体而言,消融50个高判别神经元后,文学风格指标提升25.7%。这揭示了神经网络中观察相关性与因果必要性之间的关键差距。研究挑战了‘在理想输入下激活的神经元会生成相应输出’这一假设,对机制可解释性研究和人工智能对齐具有深远影响。

原文摘要 · Abstract (English)

We present a mechanistic analysis of literary style in GPT-2, identifying individual neurons that discriminate between exemplary prose and rigid AI-generated text. Using Herman Melville's Bartleby, the Scrivener as a corpus, we extract activation patterns from 355 million parameters across 32,768 neurons in late layers. We find 27,122 statistically significant discriminative neurons ($p < 0.05$), with effect sizes up to $|d| = 1.4$. Through systematic ablation studies, we discover a paradoxical result: while these neurons correlate with literary text during analysis, removing them often improves rather than degrades generated prose quality. Specifically, ablating 50 high-discriminating neurons yields a 25.7% improvement in literary style metrics. This demonstrates a critical gap between observational correlation and causal necessity in neural networks. Our findings challenge the assumption that neurons which activate on desirable inputs will produce those outputs during generation, with implications for mechanistic interpretability research and AI alignment.

可解释性语言模型神经元分析文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。