arXiv:2605.11302cs.LGcs.AI2026-05被引 2

提出时间敏感生成理论,证明稀疏幻觉可避免模式崩溃

A Theory of Time-Sensitive Language Generation: Sparse Hallucination Beats Mode Collapse

  • 引入基于排名的生成时限机制,要求高排名文本更早输出
  • 证明强一致性生成器无法实现时间敏感生成,但稀疏幻觉可突破限制
  • 在超线性时限下可实现最优生成密度,适合追求时效性的生成场景

我们研究在全局字符串偏好排序下语言生成的极限情况,该排序由Kleinberg和Wei提出。与以往工作类似,目标是生成多样性,但新增时间性要求:高排名字符串应更早生成。字符串仅在截止前生成才被计入,其截止时间由字符串排名映射函数决定。这符合机器学习中对“更简单”或“更合理”输出的归纳偏置。我们证明,对于最终一致的生成器(多数前期工作主角),时间敏感生成在强意义上不可能实现。但在最温和的一致性松弛下——幻觉率随时间趋于零——我们可突破该不可能性。具体而言,可在任意超线性截止时间函数下实现最优生成密度。同时证明,线性截止时间下若幻觉率趋于零,则时间敏感生成仍不可行。

原文摘要 · Abstract (English)

We study language generation in the limit under a global preference ordering on strings, as introduced by Kleinberg and Wei. As is done in previous work, we aim for breadth, but impose an additional requirement of timeliness: higher-ranked strings should be generated earlier. A string is then only credited if it is generated before a deadline, where its deadline is defined by a function that maps a string's rank in the target language to the time by which it must be produced. This is in keeping with a central consideration in machine learning, where inductive bias favors ``simpler'' or ``more plausible'' outputs, all else being equal. We show that timely generation is impossible in a strong sense for eventually consistent generators -- the protagonists of most prior related work. Under what is perhaps the mildest natural relaxation of consistency, a hallucination rate that vanishes over time, we show that we can circumvent our impossibility result. In particular, we can achieve optimal density with respect to any superlinear deadline function. We also show this is tight by ruling out timely generation with linear deadlines and vanishing hallucination rate.

语言生成时间敏感幻觉控制生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。