arXiv:2504.14370math.COcs.CL2025-04被引 20

提出用密度衡量语言生成的广度,实现有效与多样性的平衡。

Density Measures for Language Generation

  • 用密度量化语言生成的覆盖范围,定义有效性与广度的权衡
  • 设计算法使生成结果在目标语言中具有严格正密度
  • 揭示生成过程需在高低密度表示间振荡以达成最优广度

大语言模型的成功引发了对语言生成理论的深入研究。近期工作提出了‘极限下的语言生成’抽象视角:生成被视为对抗者与算法之间的博弈——对抗者从未知语言 $K$ 中生成字符串($K$ 来自可数候选语言集合),算法在观察有限样本后需生成 $K$ 内但未见过的新字符串。该形式强调核心矛盾:有效性(仅生成合法语句)与广度(生成多样化输出)之间的权衡。这一权衡在实际应用中体现为幻觉与模式崩溃的平衡。尽管重要,该权衡长期缺乏定量分析。本文通过引入密度概念量化广度,发现现有算法生成集在真实语言中可能密度为零,这看似不可避免。我们证明此并非必然:提出一种算法,其输出在 $K$ 中具有严格正密度。同时分析算法内部假设语言序列,发现达到最强广度需在高、低密度表示间无限振荡。我们的分析引入语言族上的新拓扑结构,收敛与极限点概念起关键作用。

原文摘要 · Abstract (English)

The recent successes of large language models (LLMs) have led to a surge of theoretical research into language generation. A recent line of work proposes an abstract view, called language generation in the limit, where generation is seen as a game between an adversary and an algorithm: the adversary generates strings from an unknown language $K$, chosen from a countable collection of candidate languages, and after seeing a finite set of these strings, the algorithm must generate new strings from $K$ that it has not seen before. This formalism highlights a key tension: the trade-off between validity (the algorithm should only produce strings from the language) and breadth (it should be able to produce many strings from the language). This trade-off is central in applied language generation as well, where it appears as a balance between hallucination (generating invalid utterances) and mode collapse (generating only a restricted set of outputs). Despite its importance, this trade-off has been challenging to study quantitatively. We develop ways to quantify this trade-off by formalizing breadth using measures of density. Existing algorithms for language generation in the limit produce output sets that can have zero density in the true language, and this important failure of breadth might seem unavoidable. We show, however, that such a failure is not necessary: we provide an algorithm for language generation in the limit whose outputs have strictly positive density in $K$. We also study the internal representations built by these algorithms, specifically the sequence of hypothesized candidate languages they consider, and show that achieving the strongest form of breadth may require oscillating indefinitely between high- and low-density representations. Our analysis introduces a novel topology on language families, with notions of convergence and limit points playing a key role.

语言生成密度测量大模型理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。