arXiv:2607.23361cs.DScs.CL2026-07

研究语言生成中幻觉率的理论极限,发现零幻觉率也能提升生成能力。

Hallucination Rates in Language Generation

  • 引入幻觉率概念,允许算法无限生成错误但控制错误频率。
  • 证明零幻觉率下可生成有限误差无法覆盖的语言集合。
  • 揭示幻觉率与生成广度共同决定语言生成能力的严格层级结构。

Kleinberg 和 Mullainathan [KM24] 提出的生成在极限模型中,算法若在有限时间后永不犯错,则认为正确生成某语言。然而现实中,即使先进语言模型也频繁幻觉。本文首次研究带(无限)幻觉的生成在极限情形:算法可无限次生成错误,但错误发生速率受限(甚至为0测度)。我们证明,即便幻觉率为0,生成能力仍强于有限误差情形:存在有限误差无法生成、但可由无限误差生成的语言集合,即使错误仅出现在0测度时间步。尽管所有可数语言集合可通过有限误差生成,我们发现由幻觉率刻画的不可数语言集合存在严格层次结构。该层次结构延伸至生成广度(目标语言占比):所有可数集合可达最优广度1/2 [KW26b],但每个广度和幻觉率组合均存在严格分离。此外,在禁止重复字符串的设定下,我们比较正确与错误生成集的大小,再次揭示各幻觉率与广度下的严格层级。这些结果揭示了带幻觉生成在极限下的丰富结构,并确立幻觉率作为语言生成理论研究的重要参数。

原文摘要 · Abstract (English)

Language generation in the limit is an elegant model introduced by Kleinberg and Mullainathan [KM24] to formally study language generation by an algorithm that learns solely based on example strings. In this model, an algorithm is said to correctly generate from a language if it never makes an error after some finite time. In contrast, even sophisticated language models are known to regularly hallucinate in practice. In this paper, we initiate the study of language generation in the limit with (infinite) hallucination, i.e., the algorithm may generate incorrect strings infinitely often, but the errors occur at a limited rate (possibly even with 0-measure). We first show that hallucination, even at rate 0, makes generation in the limit strictly more powerful: there are language collections that cannot be generated with finite error but can be generated with infinite error, even when errors occur on a 0-measure set of time-steps. Furthermore, while all countable collections are generatable with finite error, we show a strict hierarchy of (uncountable) language collections characterized by the hallucination rate. This hierarchy extends to breadth, the fraction of the target language generated. While all countable collections can attain the optimal breadth of 1/2 [KW26b], we show strict separation at every breadth and hallucination rate. Finally, we study generation in the limit without repetition, where the algorithm may not repeat strings. This lets us compare the sets of correct and incorrect strings generated, rather than the fractions of correct and incorrect time-steps. Once again, we demonstrate a strict hierarchy at every hallucination rate and breadth. Taken together, these results reveal rich structure in language collections generatable in the limit with hallucination and establish hallucination rate as an important parameter in the theoretical study of language generation.

语言生成幻觉率理论分析计算模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。