arXiv:2504.20951cs.CL2025-04

用引力场理论解释大模型如何选词,揭示幻觉与提示敏感性的根源。

Information Gravity: A Field-Theoretic Model for Token Selection in Large Language Models

  • 将查询视为有信息质量的物体,扭曲语义空间形成引力势阱。
  • 解释幻觉源于低密度语义空洞,温度影响输出多样性。
  • 适合研究模型生成机制与提示工程的学者参考。

我们提出一种名为“信息引力”的理论模型,用于描述大语言模型(LLM)的文本生成过程。该模型借鉴场论与时空几何的物理框架,将用户查询视为具有“信息质量”的物体,通过曲率改变模型的语义空间,形成引力势阱,从而在生成过程中“吸引”相关标记。该模型为多种观测到的LLM行为提供了机制解释,包括幻觉(源于低密度语义空洞)、对查询表述的敏感性(因语义场曲率变化所致),以及采样温度对输出多样性的调节作用。

原文摘要 · Abstract (English)

We propose a theoretical model called "information gravity" to describe the text generation process in large language models (LLMs). The model uses physical apparatus from field theory and spacetime geometry to formalize the interaction between user queries and the probability distribution of generated tokens. A query is viewed as an object with "information mass" that curves the semantic space of the model, creating gravitational potential wells that "attract" tokens during generation. This model offers a mechanism to explain several observed phenomena in LLM behavior, including hallucinations (emerging from low-density semantic voids), sensitivity to query formulation (due to semantic field curvature changes), and the influence of sampling temperature on output diversity.

语言模型生成机制引力类比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。