arXiv:2601.23258cs.LGcs.AI2026-01被引 3

不依赖语言分布假设,实现更通用的语言识别与生成

Agnostic Language Identification and Generation

  • 在无前提假设下定义语言识别与生成的新目标
  • 获得接近最优的理论性能边界
  • 适合研究鲁棒性与真实场景下的语言建模

语言识别与生成的最新研究已确立了这些任务可达到的紧致统计速率。这些工作通常基于强可实现性假设:输入数据必然来自某个未知分布,且该分布支持于给定语言集合中的某一种语言。本文完全放松这一可实现性假设,不对输入数据分布施加任何限制,提出了在更通用的“非知觉”(agnostic)设置下研究语言识别与生成的新目标。在两个问题上,均获得了新颖且有趣的表征,并达到了近乎紧致的速率。

原文摘要 · Abstract (English)

Recent works on language identification and generation have established tight statistical rates at which these tasks can be achieved. These works typically operate under a strong realizability assumption: that the input data is drawn from an unknown distribution necessarily supported on some language in a given collection. In this work, we relax this assumption of realizability entirely, and impose no restrictions on the distribution of the input data. We propose objectives to study both language identification and generation in this more general "agnostic" setup. Across both problems, we obtain novel interesting characterizations and nearly tight rates.

语言识别生成模型理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。