arXiv:2511.01202cs.ITcs.AI2025-11被引 4

用语义原子TOKEN取代比特,构建大模型的语义信息理论

Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs

  • 以TOKEN为语义基本单元,重构注意力与Transformer的物理意义
  • 提出定向率失真函数,量化预训练与强化学习后训练的因果信息流
  • 揭示了语言模型推理的因果边界,适合关注模型本质原理的研究者

尽管大语言模型在实践中取得成功,但当前范式仍依赖经验与实验,严重依赖算力和数据,缺乏第一性原理理论。本文融合统计物理、信号处理与经典信息论,提出一种语义信息理论,核心是将传统无语义的比特(BIT)替换为承载意义的宏观单元TOKEN。在此框架下,注意力机制与Transformer被重释为能量模型,语义嵌入被视为语义流形上的向量表示。将大模型建模为带反馈的状态化信道,采用马塞的定向信息作为自回归生成的因果度量,推导出预训练的定向率失真函数、基于强化学习的后训练定向率-奖励函数,以及推理时语义信息流的次鞅描述。该理论精确识别了下一词预测与格兰杰因果推断的关系,并以皮尔的因果阶梯为基准,明确了大模型推理的极限——正如比特定义了信息时代,未来将由TOKEN定义人工智能时代。

原文摘要 · Abstract (English)

Despite the empirical successes of Large Language Models (LLMs), the prevailing paradigm is heuristic and experiment-driven, tethered to massive compute and data, while a first-principles theory remains absent. This treatise develops a Semantic Information Theory at the confluence of statistical physics, signal processing, and classical information theory, organized around a single paradigm shift: replacing the classical BIT - a microscopic substrate devoid of semantic content - with the macroscopic TOKEN as the atomic carrier of meaning and reasoning. Within this framework we recast attention and the Transformer as energy-based models, and interpret semantic embedding as vectorization on the semantic manifold. Modeling the LLM as a stateful channel with feedback, we adopt Massey's directed information as the native causal measure of autoregressive generation, from which we derive a *directed rate-distortion function for pre-training, a directed rate-reward function for RL-based post-training, and a sub-martingale account of inference-time semantic information flow. This machinery makes precise the identification of next-token prediction with Granger causal inference, and sharpens the limits of LLM reasoning against Pearl's Ladder of Causation - affirming that *whereas the BIT defined the Information Epoch, the TOKEN will define the AI Epoch.

大模型理论语义信息因果推理注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。