arXiv:2510.13928cs.CLcs.AI2025-10被引 3

持续接触垃圾文本会让大模型认知能力下降,研究发现其推理与安全性能显著退化。

LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X

  • 通过控制社交媒体数据的互动度和语义质量,构建垃圾与对照数据集进行实验。
  • 垃圾数据训练使模型推理、长上下文理解等能力下降,如ARC-Challenge得分从72.1降至57.2。
  • 模型无法完全恢复受损能力,提示存在持久表征漂移,需定期做认知健康检查。

我们提出并验证了大语言模型(LLM)的‘脑腐’假说:持续暴露于低质网络文本会引发大模型的长期认知衰退。为揭示垃圾内容的影响,我们在真实Twitter/X语料上设计了受控实验,通过两个正交操作维度(M1:互动程度;M2:语义质量)构建垃圾与反向对照数据集,保持词元规模与训练操作一致。相较于对照组,4个大模型在垃圾数据上持续预训练后,在推理、长上下文理解、安全性方面均出现显著下降(Hedges' g > 0.3),且“暗特质”(如自恋、反社会倾向)增强。渐进式混合垃圾与对照数据呈现剂量依赖性衰退:例如,按M1标准,当垃圾比例从0%升至100%时,ARC-Challenge(CoT)得分由72.1降至57.2,RULER-CWE从83.7降至52.3。错误溯源分析显示:推理中的思维跳步是主要损伤机制;部分但不完全可逆的修复效果表明,指令微调与清洁持续预训练虽能提升表现,但无法恢复基线能力,暗示存在持久表征漂移而非格式错配。此外,推文流行度(非语义指标)比长度更能预测脑腐效应。结果表明,数据的社会属性可能是持续预训练中模型能力退化的因果驱动因素,推动对部署中模型进行常规‘认知健康检查’。

原文摘要 · Abstract (English)

We propose and test the LLM Brain Rot Hypothesis: continual exposure to junk web text induces lasting cognitive decline in large language models (LLMs). To unveil junk effects, we designed a novel controlled experiment on real Twitter/X corpora, by constructing junk and reverse-controlled datasets via two orthogonal operationalizations: M1 (engagement degree) and M2 (semantic quality), with matched token scale and training operations across conditions. Compared to the control group, continual pre-training of 4 LLMs on the junk dataset causes non-trivial declines (Hedges' g>0.3) on reasoning, long-context understanding, safety, and inflating "dark traits" (e.g., psychopathy, narcissism). The gradual mixtures of junk and control datasets also yield dose-response cognition decay: for example, under M1, ARC-Challenge with Chain-of-Thought drops 72.1 -> 57.2 and RULER-CWE 83.7 -> 52.3 as junk ratio rises from 0% to 100%. Error forensics reveal several key insights. First, we identify thought-skipping as the primary lesion in reasoning: models increasingly truncate or skip chains. Second, partial but incomplete healing is observed: scaling instruction tuning and clean continual pre-training improve the declined cognition, yet cannot restore baseline capability, suggesting persistent representational drift rather than format mismatch. Finally, we discover that the popularity, a non-semantic metric, of a tweet is a better indicator of the Brain Rot effect than the length in M1. Together, the results provide significant, multi-perspective evidence that social effects of data could be a causal driver of LLM capability decay in continual pre-training, thereby motivating routine "cognitive health checks" for deployed and evolving LLMs.

大模型认知退化数据污染社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。