研究语言生成极限下的有效性与覆盖范围权衡,揭示了生成算法的内在限制。
Exploring Facets of Language Generation in the Limit
- 提出非均匀生成机制,确保生成正确语言样本
- 证明仅用成员查询无法实现两语言集合的非均匀生成
- 揭示生成有效性与覆盖广度不可兼得的本质矛盾
Kleinberg & Mullainathan [KM24] 提出语言生成在极限下的形式化模型:给定未知目标语言的一系列样本,目标是生成新样本,使得在某一点后不再产生错误样本。与语言识别的强负面结果相反,他们证明对所有可数语言集合,语言生成在极限下均可实现。Raman & Tewari [RT24] 研究了算法达到正确生成所需的最小输入数量,区分了统一生成(所有语言相同常数)与非统一生成(语言相关常数)。本文证明每个可数语言集合均存在具备更强非均匀生成性质的生成器。然而,尽管[KM24]的生成算法可用成员查询实现,我们证明任何算法都无法仅用成员查询对两个语言集合进行非均匀生成。此外,通过引入‘穷尽生成’定义,形式化了生成算法中有效性和覆盖范围之间的张力,并给出穷尽生成的强负面结果,表明该权衡是生成在极限中的固有属性。我们还精确刻画了穷尽生成可能的语言集合。最后,受可主动获取反馈算法启发,考虑带反馈的统一生成模型,完全以集合的复杂性测度表征此类生成可能的语言集合。
原文摘要 · Abstract (English)
The recent work of Kleinberg & Mullainathan [KM24] provides a concrete model for language generation in the limit: given a sequence of examples from an unknown target language, the goal is to generate new examples from the target language such that no incorrect examples are generated beyond some point. In sharp contrast to strong negative results for the closely related problem of language identification, they establish positive results for language generation in the limit for all countable collections of languages. Follow-up work by Raman & Tewari [RT24] studies bounds on the number of distinct inputs required by an algorithm before correct language generation is achieved -- namely, whether this is a constant for all languages in the collection (uniform generation) or a language-dependent constant (non-uniform generation). We show that every countable language collection has a generator which has the stronger property of non-uniform generation in the limit. However, while the generation algorithm of [KM24] can be implemented using membership queries, we show that any algorithm cannot non-uniformly generate even for collections of just two languages, using only membership queries. We also formalize the tension between validity and breadth in the generation algorithm of [KM24] by introducing a definition of exhaustive generation, and show a strong negative result for exhaustive generation. Our result shows that a tradeoff between validity and breadth is inherent for generation in the limit. We also provide a precise characterization of the language collections for which exhaustive generation is possible. Finally, inspired by algorithms that can choose to obtain feedback, we consider a model of uniform generation with feedback, completely characterizing language collections for which such uniform generation with feedback is possible in terms of a complexity measure of the collection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。