arXiv:2606.28354cs.CLcs.FL2026-06被引 1

允许无限次错误但频率趋零,提升语言生成的覆盖率。

Generating in the Limit with Infinitely Many Hallucinations

  • 放宽验证限制:允许无限次生成错误,只要错误频率趋近于零。
  • 在对手永久隐藏大部分语言时,召回率显著提升。
  • 适合研究大模型生成容错机制与真实场景下的生成策略。

语言生成在极限框架下将学习目标从识别未知语言转向生成该语言中未见过的有效字符串。现有研究揭示了覆盖范围与生成有效性之间的根本矛盾:广泛覆盖常以牺牲正确性为代价。本文提出新的精确度定义,将问题重构为经典的召回-精确度权衡。通过分析枚举、新颖性和有效性约束下的生成行为,我们研究了非最终有效的学习者——允许无限次错误,但要求错误频率趋于零,从而保持精确度为1。结果表明,在对手持续隐藏大量目标语言的情况下,此松弛可严格提高召回率。此外,我们还研究了新颖性约束的连续松弛,仅要求固定比例输出为新内容。整体上,这些发现推动了更贴近大型语言模型实际生成场景的理论建模,承认错误和重复不可避免,但其发生率可控。

原文摘要 · Abstract (English)

The classic paradigm of language identification in the limit models learning as a game between an adversary, who reveals strings from an unknown target language, and a learner tasked with identifying that language. The recently introduced framework of language generation in the limit shifted the objective to better reflect modern language modeling, requiring the learner to produce valid, unseen strings from the target language. Related work highlighted a fundamental tension: a broad coverage of the target often comes at the cost of validity. We introduce a new notion of precision and recast this problem as the classic recall-precision trade-off. We analyze generation in the limit under varying constraints on enumeration, novelty, and validity, aimed at reflecting settings closer to those encountered by large language models. A key contribution is our analysis of learners that are not eventually valid: we allow infinitely many mistakes, provided their frequency tends to zero so that precision remains one. We show that this relaxation can strictly increase recall when the adversary permanently withholds a large portion of the target language. We also study a continuous relaxation of the novelty constraint that requires only a fixed fraction of outputs to be novel. Taken together, our results move toward a more realistic model of language generation where occasional errors and repetitions are unavoidable, but their rates are controlled.

语言生成生成模型理论分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。